【问题标题】:Python regex findallPython 正则表达式 findall
【发布时间】:2011-12-06 20:13:38
【问题描述】:

我正在尝试使用 Python 2.7.2 中的正则表达式从字符串中提取所有出现的标记词。或者简单地说,我想提取[p][/p] 标签内的每一段文本。 这是我的尝试:

regex = ur"[\u005B1P\u005D.+?\u005B\u002FP\u005D]+?"
line = "President [P] Barack Obama [/P] met Microsoft founder [P] Bill Gates [/P], yesterday."
person = re.findall(pattern, line)

打印person 产生['President [P]', '[/P]', '[P] Bill Gates [/P]']

正确的正则表达式是什么:['[P] Barack Obama [/P]', '[P] Bill Gates [/p]'] 或['Barrack Obama', 'Bill Gates']。

【问题讨论】:

    标签: python regex


    【解决方案1】:
    import re
    regex = ur"\[P\] (.+?) \[/P\]+?"
    line = "President [P] Barack Obama [/P] met Microsoft founder [P] Bill Gates [/P], yesterday."
    person = re.findall(regex, line)
    print(person)
    

    产量

    ['Barack Obama', 'Bill Gates']
    

    正则表达式ur"[\u005B1P\u005D.+?\u005B\u002FP\u005D]+?" 完全相同 unicode 为 u'[[1P].+?[/P]]+?',但更难阅读。

    第一个括号组[[1P] 告诉re 列表中的任何字符['[', '1', 'P'] 应该匹配,与第二个括号组[/P]] 类似。这根本不是你想要的。所以,

    • 删除外部封闭方括号。 (同时删除 在P前面流浪1。)
    • 要保护[P] 中的文字括号,请使用a 转义括号 反斜杠:\[P\].
    • 要仅返回标签内的单词,请放置分组括号 .+? 附近。

    【讨论】:

      【解决方案2】:

      试试这个:

         for match in re.finditer(r"\[P[^\]]*\](.*?)\[/P\]", subject):
              # match start: match.start()
              # match end (exclusive): match.end()
              # matched text: match.group()
      

      【讨论】:

      • 我真的很喜欢这个答案。如果您只想处理匹配项,则无需任何额外语句,例如 1) 保存列表,2) 处理列表不等同于 str = 'purple alice@google.com, blah monkey bob@abc.com blah diswasher' ## 这里 re.findall() 返回所有找到的电子邮件字符串的列表 emails = re.findall(r'[\w\.-]+@[\w\.-]+', str) # # ['alice@google.com', 'bob@abc.com'] for email in emails: # 对找到的每个电子邮件字符串执行某些操作 print email
      【解决方案3】:

      您的问题不是 100% 清楚,但我假设您想找到 [P][/P] 标签中的每一段文字:

      >>> import re
      >>> line = "President [P] Barack Obama [/P] met Microsoft founder [P] Bill Gates [/P], yesterday."
      >>> re.findall('\[P\]\s?(.+?)\s?\[\/P\]', line)
      ['Barack Obama', 'Bill Gates']
      

      【讨论】:

        【解决方案4】:

        你可以用

        替换你的模式
        regex = ur"\[P\]([\w\s]+)\[\/P\]"
        

        【讨论】:

        • 注意格式; 使用预览区域。因为你没有正确格式化它,所以反斜杠被浪费了(markdown 就是这样糟糕)。
        • 你为什么用[\w\s]+而不是他用的.*??无论如何,在我看来.*? 更有可能是他想要的。 [\w\s] 非常有限。
        • 故意的限制。我使用 [\w\s]+ 因为显然提问者想要提取很少包含数字的名称。另请注意,提问者想要提取单词,而不是数字。只是我的意见,cmiiw
        • 那些带有重音等有趣特征的名字呢? not re.match('\w', u'é')。如果名称是任意的,您不应忽视非拉丁名称的可能性。
        【解决方案5】:

        使用这种模式,

        pattern = '\[P\].+?\[\/P\]'

        查看here

        【讨论】:

          猜你喜欢
          • 2015-08-13
          • 1970-01-01
          • 1970-01-01
          • 2011-07-18
          • 1970-01-01
          • 2013-06-30
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多