【问题标题】:Matching multiple words conditionally有条件地匹配多个单词
【发布时间】:2019-09-22 23:51:44
【问题描述】:

我有一个大文本文件,其中包含类似于以下行的行:

timestamp = foo bar baz
timestamp = foo bar
timestamp = foo

我试图编写一个匹配 foo 的正则表达式,但如果 bar 和 baz 都存在,它也匹配那些。

r"= (.*) (.*)? (.*)?"

但它只匹配foo bar baz 字符串,而不匹配其他两个。如何使正则表达式与可选项匹配?

【问题讨论】:

  • 像= *(\S+)(?: *(\S+))?(?: *(\S+))? 这样的东西对你有用。见the regex demo。
  • 唯一的小问题是它与 foo bar 案例的 = 匹配 - 不确定为什么:)
  • @CaseyJones 添加积极的后视 (?
  • if line.endwith('foo') ... ???
  • @CaseyJones 匹配= 有什么问题?我知道您只对捕获的子字符串感兴趣。

标签: python regex


【解决方案1】:

我猜你可能会使用一些简单的表达式来获得所需的输出,例如:

(\w+\s*=\s*)|(\w+)

测试

import re


regex = r"(\w+\s*=\s*)|(\w+)"
string = """
timestamp = foo bar baz foo bar baz
timestamp = foo bar baz
timestamp = foo bar
timestamp = foo
"""

for groups in re.findall(regex, string):
    if groups[0] == '':
        print(groups[1])
    else:
        print("--- next timestamp ----")

输出

--- next timestamp ----
foo
bar
baz
foo
bar
baz
--- next timestamp ----
foo
bar
baz
--- next timestamp ----
foo
bar
--- next timestamp ----
foo

如果您希望简化/修改/探索表达式,在regex101.com 的右上角面板中已对此进行了说明。如果您愿意,您还可以在this link 中观看它如何与一些示例输入匹配。


【讨论】:

    【解决方案2】:

    也许这样就足够了?

     (?<=\=\s)(\S+)\s?(\S+)? ?(\S+)?
    

    Regex Demo

    解释:

     (?<=\=\s)       # Positive lookbehind - capture = + space but don't match
     (\S+)           # Capture any non-whitespace character
     \s?             # Capture optional space
     (\S+)?          # Capture any non-whitespace character
      ?              # Capture optional space
     (\S+)?          # Capture any non-whitespace character
    

    【讨论】:

      【解决方案3】:

      你可以使用

      r'= *(\S+)(?: *(\S+))?(?: *(\S+))?'
      

      或者,匹配任何水平空格:

      r'=[^\S\r\n]*(\S+)(?:[^\S\r\n]*(\S+))?(?:[^\S\r\n]*(\S+))?'
      

      见regex demo

      详情

      • =[^\S\r\n]* - 一个 = 字符和除 LF、CR 和非空格(即,除换行符和回车之外的所有空格)之外的任何 0 个或多个字符,或者如果您使用 *,则只是空格
      • (\S+) - 第 1 组:任何 1+ 个非空白字符
      • (?:[^\S\r\n]*(\S+))? - 一个可选的非捕获组,匹配 1 次或 0 次
        • [^\S\r\n]* - 0+ 个水平空格
        • (\S+) - 第 2 组:任何 1+ 个非空白字符
      • (?:[^\S\r\n]*(\S+))? - 一个可选的非捕获组,匹配 1 次或 0 次
        • [^\S\r\n]* - 0+ 个水平空格
        • (\S+) - 第 3 组:任何 1+ 个非空白字符

      Python demo:

      import re
      s = "timestamp = foo bar baz\ntimestamp = foo bar\ntimestamp = foo"
      print( re.findall(r'= *(\S+)(?: *(\S+))?(?: *(\S+))?', s) )
      # => [('foo', 'bar', 'baz'), ('foo', 'bar', ''), ('foo', '', '')]
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2013-12-08
        • 2022-01-11
        • 1970-01-01
        • 1970-01-01
        • 2011-12-21
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多