【问题标题】:Replace strings in a pandas column替换熊猫列中的字符串
【发布时间】:2019-11-23 17:14:36
【问题描述】:

背景

我有以下示例 df,其中包含 PHYSICIAN 在 Text 列中,后跟医生姓名(以下所有名称都是虚构的)

import pandas as pd
df = pd.DataFrame({'Text' : ['PHYSICIAN: Jon J Smith was here today', 
                                   'And Mary Lisa Rider found here', 
                                   'Her PHYSICIAN: Jane A Doe is also here',
                                ' She was seen by  PHYSICIAN: Tom Tucker '], 

                      'P_ID': [1,2,3,4],
                      'N_ID' : ['A1', 'A2', 'A3', 'A4']

                     })

#rearrange columns
df = df[['Text','N_ID', 'P_ID']]
df

                                     Text         N_ID  P_ID
0   PHYSICIAN: Jon J Smith was here today           A1  1
1   And Mary Lisa Rider found here                  A2  2
2   Her PHYSICIAN: Jane A Doe is also here          A3  3
3   She was seen by PHYSICIAN: Tom Tucker           A4  4

目标

1) 将单词PHYSICIAN 后面的名称(例如PHYSICIAN: Jon J Smith)替换为PHYSICIAN: **BLOCK**

2) 创建一个名为Text_Phys的新列

期望的输出

                                  Text            N_ID P_ID  Text_Phys
0   PHYSICIAN: Jon J Smith was here today           A1  1   PHYSICIAN: **BLOCK** was here today
1   And Mary Lisa Rider found here                  A2  2   And Mary Lisa Rider found here
2   Her PHYSICIAN: Jane A Doe is also here          A3  3   Her PHYSICIAN: **BLOCK** is also here
3   She was seen by PHYSICIAN: Tom Tucker           A4  4   She was seen by PHYSICIAN: **BLOCK**

我已经尝试了以下

1) df['Text_Phys'] = df['Text'].replace(r'ABC.*', 'ABC: ***BLOCK***', regex=True)

2)df['Text_Phys'] = df['Text'].replace(r'ABC\s+', 'ABC: ***BLOCK***', regex=True)

但它们似乎不太好用

问题

如何实现我想要的输出?

【问题讨论】:

  • 它应该作为df['Text'] = df['Text'].replace(r'PHYSICIAN', 'PHYSICIAN: ***PHI***', regex=True)和df['Text'] = df['Text'].replace(r'Physician', 'Physician: ***PHI***', regex=True)工作
  • 我试过了,但不太奏效
  • import re 然后df['Text_Phys'] = df['Text'].str.replace('PHYSICIAN', 'PHYSICIAN: ***PHI***', flags=re.I) 怎么样,但它会使情况在上层。但是,早期的工作对我来说很好,你使用的是什么版本的熊猫。
  • 如何识别文本的哪一部分是医师姓名?
  • PHYSICIAN: 后面的子串部分很容易得到。但是,几乎不可能在子字符串中识别Jon J Smith、Jane A Doe 和Tom Tucker。除非您有一些规则来识别它们,否则您怎么知道它们是要替换的名称?

标签: python python-3.x string pandas text


【解决方案1】:

试试这个:使用正则表达式来定义你想匹配的单词和位置 您想停止搜索(您可以生成所有单词的列表 在“**”之后发生以进一步自动化代码)。而不是 为了时间,我做了“Found|was |is”的快速硬编码。

代码如下:

import pandas as pd
df = pd.DataFrame({'Text' : ['PHYSICIAN: Jon J Smith was here today', 
                                   'And his Physician: Mary Lisa Rider found here', 
                                   'Her PHYSICIAN: Jane A Doe is also here',
                                ' She was seen by  PHYSICIAN: Tom Tucker '], 

                      'P_ID': [1,2,3,4],
                      'N_ID' : ['A1', 'A2', 'A3', 'A4']

                     })

df = df[['Text','N_ID', 'P_ID']]
df
    Text    N_ID    P_ID
0   PHYSICIAN: Jon J Smith was here today   A1  1
1   And his Physician: Mary Lisa Rider found here   A2  2
2   Her PHYSICIAN: Jane A Doe is also here  A3  3
3   She was seen by PHYSICIAN: Tom Tucker   A4  4

word_before = r'PHYSICIAN:'
words_after = r'.*?(?=found |was |is )'
words_all =r'PHYSICIAN:[\w\s]+'

import re

pattern = re.compile(word_before+words_after, re.IGNORECASE)
pattern2 = re.compile(words_all, re.IGNORECASE)

for i in range(len(df['Text'])):
    df.iloc[i,0] = re.sub(pattern,"PHYSICIAN: **BLOCK** ", df["Text"][i])
    if 'PHYSICIAN: **BLOCK**' not in df.iloc[i,0]:
        df.iloc[i,0] = re.sub(pattern2,"PHYSICIAN: **BLOCK** ", df["Text"][i])

df
    Text    N_ID    P_ID
0   PHYSICIAN: **BLOCK** was here today A1  1
1   And his PHYSICIAN: **BLOCK** found here A2  2
2   Her PHYSICIAN: **BLOCK** is also here   A3  3
3   She was seen by PHYSICIAN: **BLOCK**    A4  4

【讨论】:

    猜你喜欢
    • 2019-06-03
    • 2017-03-12
    • 2019-02-06
    • 2021-05-15
    • 2020-03-07
    • 2020-04-08
    • 2019-03-04
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多