【问题标题】:Regex Python [python-2.7]正则表达式 Python [python-2.7]
【发布时间】:2015-09-22 07:26:53
【问题描述】:

我正在开发一个 Python 程序,该程序可以筛选 .txt 文件以查找属名和种名。这些行的格式如下(是的,等号始终围绕通用名称):

1. =Common Name= Genus Species some other words that I don't want.
2. =Common Name= Genus Species some other words that I don't want.

我似乎无法找出一个只匹配属和种而不匹配通用名称的正则表达式。我知道等号 (=) 可能会有所帮助,但我想不出如何使用它们。

编辑:一些真实数据:

1. =Western grebe.= ÆCHMOPHORUS OCCIDENTALIS. Rare migrant; western species, chiefly interior regions of North America.

2. =Holboell's grebe.= COLYMBUS HOLBOELLII. Rare migrant; breeds far north; range, all of North America.

3. =Horned grebe.= COLYMBUS AURITUS. Rare migrant; range, almost the same as the last.

4. =American eared grebe.= COLYMBUS NIGRICOLLIS CALIFORNICUS. Summer resident; rare in eastern, common in western Colorado; breeds from plains to 8,000 feet; partial to alkali lakes; western species.

【问题讨论】:

  • 你想要什么作为你的例子的输出?
  • 属物种(起点、终点)
  • 你能告诉我们你解决这个问题的尝试吗?
  • 你能给我们看一些真实的输入数据吗?
  • 抱歉,我花了这么长时间才回复一些真实的输入数据:1. =Western grebe.= ÆCHMOPHORUS OCCIDENTALIS。稀有移民;西部物种,主要是北美的内陆地区。 2. =Holboell's grebe.= COLYMBUS HOLBOELLII.稀有移民;繁殖遥远的北方;范围,整个北美。 3. =角鸊鷉。= COLYMBUS AURITUS。稀有移民;范围,几乎和上次一样。 4. =美耳鸊鷉。= COLYMBUS NIGRICOLLIS CALIFORNICUS。夏季居民;东部少见,科罗拉多西部常见;从平原到 8,000 英尺的地方繁殖;偏碱湖;西方物种。

标签: python regex python-2.7


【解决方案1】:

你可能不需要这个正则表达式。如果您需要的单词的顺序和单词的数量始终相同,您可以将每一行拆分为子字符串列表并获取该列表的第三个(属)和第四个(种)元素。代码可能如下所示:

myfile = open('myfilename.txt', 'r')
for line in myfile.readlines():
    words = line.split()
    genus, species = words[2], words[3]

对我来说,它看起来更像“pythonic”。

如果通用名称可以包含多个单词,则建议的代码将返回错误的结果。要在这种情况下也获得正确的结果,您可以使用以下代码:

myfile = open('myfilename.txt', 'r')
for line in myfile.readlines():
    words = line.split('=')[2].split() # If the program returns wrong results, try changing the index from 2 to 1 or 3. What number is the right one depends on whether there can be any symbols before the first "=".
    genus, species = words[0], words[1]

【讨论】:

  • 同意。正则表达式对此太过分了。
  • 但我认为通用名称可能不止一个单词(我假设这是关于生物物种),但是首先使用 = 作为分隔符,然后按单词拆分可以使用这样的输入跨度>
  • @m.cekiera 是的,我没想到,我会编辑我的答案,谢谢。
【解决方案2】:

如果足以按组捕获单词(并且您不会直接匹配),您可以尝试:

(?=\d\.\s*=[^=]+=\s(?:(?P<genus>\w+)\s(?P<species>\w+)))

DEMO

所需的值将位于 &lt;genus&gt; 和 &lt;species&gt; 组中。整个正则表达式是正则表达式,因此它匹配字符串开头的零点位置,但它会将一些内容捕获到组中。

  • (?=\d\.\s*=[^=]+=\s - 小数后跟相等的一些内容 标志和空间,
  • (?:(?P&lt;genus&gt;\w+)\s(?P&lt;species&gt;\w+))) - 捕获第一个词到属 组,第二个词是物种组,

【讨论】:

    【解决方案3】:

    你可以试试这样的:

    import re
    
    txt='1. =Common Name= Genus Species some other words that I don\'t want.'
    
    re1='.*?'   # Non-greedy match on filler
    re2='(?:[a-z][a-z]+)'   # Uninteresting: word
    re3='.*?'   # Non-greedy match on filler
    re4='(?:[a-z][a-z]+)'   # Uninteresting: word
    re5='.*?'   # Non-greedy match on filler
    re6='((?:[a-z][a-z]+))' # Word 1
    re7='.*?'   # Non-greedy match on filler
    re8='((?:[a-z][a-z]+))' # Word 2
    
    rg = re.compile(re1+re2+re3+re4+re5+re6+re7+re8,re.IGNORECASE|re.DOTALL)
    m = rg.search(txt)
    if m:
        word1=m.group(1)
        word2=m.group(2)
        print "("+word1+")"+"("+word2+")"+"\n"
    

    在您的测试输入中,如 txt 所示,这将打印出来

    (属)(种)

    你可以this 很棒的网站来帮助做这样的正则表达式!

    希望对你有帮助

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2012-07-06
      • 1970-01-01
      • 1970-01-01
      • 2011-12-29
      • 2016-11-29
      • 2017-04-10
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多