【问题标题】:Extract words based on regex基于正则表达式提取单词
【发布时间】:2021-12-24 06:47:15
【问题描述】:

我正在尝试提取单词 SUITE、suite、ste、Ste 之后出现的任何内容。或熊猫数据框列中的字符#。示例如下

4230 Harding Pike Ste 435
4230 Harding Pike Suite 435
4230 Harding Pike SUITE A
4230 Harding Pike Ste. 101
4230 Harding Pike SUITE B-200
4230 Harding Pike #900
4230 Harding Pike STE 503
4230 Harding Pike SUITE 300
4230 Harding Pike Ste 700
4230 Harding Pike SUITE #2

结果:

SuiteNos

435
435  
A
101
B-200
900
503
300
700
2

我尝试了以下正则表达式,但它没有按预期工作 -

 df_merged['address'].str.extract(r'\**\**suite.*\**\**')
 df_merged['address'] = re.findall('@([suite]+)', df_merged['address'])

【问题讨论】:

    标签: python regex pandas


    【解决方案1】:

    使用不区分大小写的匹配,您可以使用带有交替的捕获组:

    (?i)(?:#|Suite|ste\.?)\s*([^\s#].*)
    

    模式匹配:

    • (?:非捕获组
      • Suite 字面匹配
      • |或者
      • ste\.? 将 ste 与可选的点匹配
      • |或者
      • # 字面匹配
    • )关闭非捕获组
    • \s* 匹配可选的空白字符
    • ([^\s#].*) 捕获组 1,从匹配不带 # 的非空白字符和该行的其余部分开始

    Regex demo

    import pandas as pd
    
    strings = [
        "4230 Harding Pike Ste 435",
        "4230 Harding Pike Suite 435",
        "4230 Harding Pike SUITE A",
        "4230 Harding Pike Ste. 101",
        "4230 Harding Pike SUITE B-200",
        "4230 Harding Pike #900",
        "4230 Harding Pike STE 503",
        "4230 Harding Pike SUITE 300",
        "4230 Harding Pike Ste 700",
        "4230 Harding Pike SUITE #2"
    ]
    
    pattern = r"(?i)(?:#|Suite|ste\.?)\s*([^\s#].*)"
    df_merged = pd.DataFrame(strings, columns = ['address'])
    df_merged['SuiteNos'] = df_merged['address'].str.extract(pattern)
    
    print(df_merged["SuiteNos"])
    

    输出

    0      435
    1      435
    2        A
    3      101
    4    B-200
    5      900
    6      503
    7      300
    8      700
    9        2
    

    【讨论】:

    • 例子表明#后面只能跟数字。如果是这样,您将需要一个小模块。顺便说一句,这是第 100 万个仅通过示例陈述问题而产生歧义的示例。
    【解决方案2】:

    使用您显示的示例,请尝试以下正则表达式。

    (?:(?:[Ss](?:[uU][iI])?[tT][eE]\.?\s+)#?|#)(\S+)
    

    Pandas 中运行以下代码:

    df['value'].str.extract(r'(?:(?:[Ss](?:[uU][iI])?[tT][eE]\.?\s+)#?|#)(\S+)')
    

    以上代码在 Pandas 中的输出以及 OP 提供的示例如下:

    0      435
    1      435
    2        A
    3      101
    4    B-200
    5      NaN
    6      503
    7      300
    8      700
    9        2
    Name: value, dtype: object
    

    Online demo for above regex

    说明:为上述添加详细说明。

    (?:                   ##Starting a non-capturing group here.
      (?:                 ##Starting another non-capturing group here.
         [Ss]             ##Matching small OR capital S here.
         (?:[uU][iI])?    ##In a non-capturing group matching u/U i/I and keep this optional.
         [tT][eE]\.?\s+   ##Matching t or T followed by e or E here followed by optional dot and 1 or more spaces here.
      )                   ##Closing 2nd non-capturing group here.
      #?|#                ##Matching # keeping it optional OR matching it no optional.
    )                     ##Closing 1st non-capturing group here.
    (\S+)                 ##Creating 1st capturing group which contains all non-spaces in it.
    

    感谢@The Fourth Bird 帮助进行调整,使仅 1 个捕获组捕获所有匹配的东西。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2011-12-31
      • 1970-01-01
      • 1970-01-01
      • 2019-12-31
      • 2021-10-30
      • 1970-01-01
      • 2016-12-21
      相关资源
      最近更新 更多