【问题标题】:Multi-pattern nested regex using Python re使用 Python re 的多模式嵌套正则表达式
【发布时间】:2020-06-17 10:12:21
【问题描述】:

使用 re 库在 Python 中编写这个 sed 程序的等效方法是什么?这种 sed 模式一次性完成搜索,而且效率很高。我正在尝试提取 cpu 的型号。请在底部查看我的 Python 代码尝试。

示例输入:

processor       : 0
vendor_id       : GenuineIntel
cpu family      : 6
model           : 45
model name      : Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz
stepping        : 6

输出:

E5-2660

示例输入 2:

processor       : 127
vendor_id       : AuthenticAMD
cpu family      : 23
model           : 1
model name      : AMD EPYC 7601 32-Core Processor
stepping        : 2

输出:

EPYC 7601

Sed:

/AuthenticAMD/{
    s/.*/AMD/p
}
/GenuineIntel/ {
    n
    n
    n
    /Celeron/ {            
        s/.*\([egptEGPT][1-9][0-9][0-9][0-9][a-zA-Z][a-zA-Z]\).*/\1/p 
        s/.*\([egptEGPT][1-9][0-9][0-9][0-9][a-zA-Z]\).*/\1/p 
        s/.*\([egptEGPT][1-9][0-9][0-9][0-9]\).*/\1/p
        q
    }
    /Xeon/ {
        s/.*[eE][3579]-\([1-9][1-9][1-9][1-9]\).*/\1/p
        s/.*\([eElL]C[1-9][0-9][0-9][0-9]\).*/\1/p
        s/.*\([35][0-9][0-9][0-9]\).*/\1/p
        q
    }
}

在 Python 中尝试(不工作):

我的代码搜索每个表达式并且不遵循任何嵌套规则,效率不高。寻找一种更好的方式来写这个。

string = """processor       : 0
            vendor_id       : GenuineIntel
            cpu family      : 6
            model           : 45
            model name      : Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz
            stepping        : 6"""

pattern = r'''GenuineIntel.*
                (?=Celeron
                    .*([egptEGPT][1-9][0-9][0-9][0-9][a-zA-Z][a-zA-Z]).*
                    .*([egptEGPT][1-9][0-9][0-9][0-9][a-zA-Z]).*
                    .*([egptEGPT][1-9][0-9][0-9][0-9]).*)|
                (?=Xeon
                    .*([eE][3579]-[1-9][0-9][0-9][0-9]).*)'''

print(re.search(pattern, string, re.MULTILINE|re.DOTALL|re.VERBOSE).groups())

【问题讨论】:

  • 对我来说看起来没问题,只要它有效。也许将 [0-9][0-9][0-9] 写成 \d{3} 并使用忽略大小写标志
  • 在您的sed 中,/AuthenticAMC/{ 不应该是/AuthenticAMD/{ 吗?另外,[eE][3579]-\([1-9][1-9][1-9][1-9]\) 中的[1-9]s 呢?你确定不是[eE][3579]-\([1-9][0-9][0-9][0-9]\)?
  • 另外,我怀疑您的sed 命令是否有效,请参阅ideone.com/NSIPMF。你说输出应该是EPYC 7601,但是 - 根据sed 命令逻辑 - 输出是AMD。
  • 感谢您更新问题的反馈。虽然我的主要目标是能够在 python 中编写正则表达式,以便它检查“Celeron”是否存在,如果不存在,那么它会跳过在 Celeron 中搜索正则表达式并移动到下一个“Xeon”。目前,我认为我编写的 python 代码会检查每个嵌套表达式?我错了吗?
  • 目前,正则表达式要求Celeron 和Xeon 出现在GenuineIntel 的右侧。我在 Python 中为 sed 命令编写了一个直接翻译,但它并没有返回你所期望的。

标签: python regex sed


【解决方案1】:

拥有像 Python 这样功能齐全的语言和结构良好的数据,我不会尝试使用正则表达式解析所有内容。相反,我只是写了一个代码来完成这项工作,只在最后使用正则表达式。通过这种方式,我可以使用非常简单的正则表达式获得简短易读的代码,而不是庞大的正则表达式。

data = {}
for line in string.split("\n"):
    left, right = line.split(":")
    data[left.strip()] = right.strip()

if data["vendor_id"] == "GenuineIntel":
    model = data["model name"]
    if "Xeon" in model:
        code = re.search(r"\bE\d-\d{4}\b", model, re.I).group(0)
    elif "Celeron" in model:
        code = re.search(r"\b[EGPT]\d{4}[a-z]{0,2}\b", model, re.I).group(0)

print(code)

关于效率——只要你没有数以百万计的字符串要解析,你就不必担心。

【讨论】:

  • 我喜欢您在输入字符串中拆分字段的方式以及您如何将正则表达式语句压缩为一个。谢谢!
【解决方案2】:

牢记https://blog.codinghorror.com/regular-expressions-now-you-have-two-problems/ 的好建议,不要将你的 python 代码基于 sed 脚本,从像这个 awk 脚本这样简单的好东西开始:

$ cat tst.awk
/^vendor_id/ {
    vendor = $NF
}
/^model name/ {
    model = "Unknown"
    if ( vendor == "GenuineIntel" ) {
        model = $7
    }
    else if ( vendor == "AuthenticAMD" ) {
        model = $5 " " $6
    }
    print model
}

我相信您也可以在 python 中轻松实现。这是一个在一个文件中使用 2 个样本输入块的示例:

$ cat file
processor       : 0
vendor_id       : GenuineIntel
cpu family      : 6
model           : 45
model name      : Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz
stepping        : 6

processor       : 127
vendor_id       : AuthenticAMD
cpu family      : 23
model           : 1
model name      : AMD EPYC 7601 32-Core Processor
stepping        : 2

$ awk -f tst.awk file
E5-2660
EPYC 7601

【讨论】:

    猜你喜欢
    • 2023-02-07
    • 1970-01-01
    • 1970-01-01
    • 2012-05-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多