【发布时间】:2019-11-23 17:14:36
【问题描述】:
背景
我有以下示例 df,其中包含 PHYSICIAN 在 Text 列中,后跟医生姓名(以下所有名称都是虚构的)
import pandas as pd
df = pd.DataFrame({'Text' : ['PHYSICIAN: Jon J Smith was here today',
'And Mary Lisa Rider found here',
'Her PHYSICIAN: Jane A Doe is also here',
' She was seen by PHYSICIAN: Tom Tucker '],
'P_ID': [1,2,3,4],
'N_ID' : ['A1', 'A2', 'A3', 'A4']
})
#rearrange columns
df = df[['Text','N_ID', 'P_ID']]
df
Text N_ID P_ID
0 PHYSICIAN: Jon J Smith was here today A1 1
1 And Mary Lisa Rider found here A2 2
2 Her PHYSICIAN: Jane A Doe is also here A3 3
3 She was seen by PHYSICIAN: Tom Tucker A4 4
目标
1) 将单词PHYSICIAN 后面的名称(例如PHYSICIAN: Jon J Smith)替换为PHYSICIAN: **BLOCK**
2) 创建一个名为Text_Phys的新列
期望的输出
Text N_ID P_ID Text_Phys
0 PHYSICIAN: Jon J Smith was here today A1 1 PHYSICIAN: **BLOCK** was here today
1 And Mary Lisa Rider found here A2 2 And Mary Lisa Rider found here
2 Her PHYSICIAN: Jane A Doe is also here A3 3 Her PHYSICIAN: **BLOCK** is also here
3 She was seen by PHYSICIAN: Tom Tucker A4 4 She was seen by PHYSICIAN: **BLOCK**
我已经尝试了以下
1) df['Text_Phys'] = df['Text'].replace(r'ABC.*', 'ABC: ***BLOCK***', regex=True)
2)df['Text_Phys'] = df['Text'].replace(r'ABC\s+', 'ABC: ***BLOCK***', regex=True)
但它们似乎不太好用
问题
如何实现我想要的输出?
【问题讨论】:
-
它应该作为
df['Text'] = df['Text'].replace(r'PHYSICIAN', 'PHYSICIAN: ***PHI***', regex=True)和df['Text'] = df['Text'].replace(r'Physician', 'Physician: ***PHI***', regex=True)工作 -
我试过了,但不太奏效
-
import re然后df['Text_Phys'] = df['Text'].str.replace('PHYSICIAN', 'PHYSICIAN: ***PHI***', flags=re.I)怎么样,但它会使情况在上层。但是,早期的工作对我来说很好,你使用的是什么版本的熊猫。 -
如何识别文本的哪一部分是医师姓名?
-
PHYSICIAN:后面的子串部分很容易得到。但是,几乎不可能在子字符串中识别Jon J Smith、Jane A Doe和Tom Tucker。除非您有一些规则来识别它们,否则您怎么知道它们是要替换的名称?
标签: python python-3.x string pandas text