【问题标题】:Extracting a specific word using Regex in Pandas在 Pandas 中使用正则表达式提取特定单词
【发布时间】:2021-02-26 01:03:50
【问题描述】:

我正在尝试从以下数据框中提取国家/地区名称

country
0   NaN
1   Country: America
2   Country: France ...More CountriesFranceNorwayP...
3   NaN
4   Country: India

使用下面的正则表达式

import re
regex = re.compile(\
    r"Country: (?P<country>\w+)"
    )

df['country'] = df['country'].str.extractall(regex).droplevel(1)

但它会返回

country
0   NaN
1   NaN
2   NaN
3   NaN
4   NaN

而不是返回

country
0   NaN
1   America
2   France
3   NaN
4   India

我错过了什么?

请指教

【问题讨论】:

    标签: python regex pandas


    【解决方案1】:

    你可以使用extract:

    df['country'] = df['country'].str.extract(r'Country:\s*(\w+)')
    

    熊猫测试:

    import pandas as pd
    import numpy as np
    df = pd.DataFrame({'country' : [np.nan, 'Country: America', 'Country France ... More countries...']})
    df['country'].str.extract(r'Country:\s*(\w+)')
    #          0
    # 0      NaN
    # 1  America
    # 2      NaN
    

    【讨论】:

    • 效果很好!我的语句返回 NaN 的任何原因?
    • @Luke 我不确定可能是什么原因,我没有看到全部数据。
    【解决方案2】:

    你也可以避开regex而使用Series.str.split:

    In [86]: df = pd.DataFrame({'country' : [np.nan, 'Country: America', 'Country: France ... More countries...', np.nan, 'Country: India']})
    
    In [87]: df
    Out[87]: 
                                     country
    0                                    NaN
    1                       Country: America
    2  Country: France ... More countries...
    3                                    NaN
    4                         Country: India
    
    In [94]: df.country.str.split(':').str[1].str.split().str[0]
    Out[94]: 
    0        NaN
    1    America
    2     France
    3        NaN
    4      India
    Name: country, dtype: object
    

    【讨论】:

    • 绝对试过了!和正则表达式一样好用!但我不确定它是否一直有效
    • 答案中的假设是基于您的样本数据的方式。可能需要对您的完整数据进行一些调整。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-12-31
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多