【问题标题】:Python regex fullmatch doesn't work as expectedPython regex fullmatch 无法按预期工作
【发布时间】:2021-11-23 02:07:03
【问题描述】:

我有一个包含一些句子的文本文件,我正在检查它们是否是基于某些规则的有效句子,并将有效或无效写入单独的文本文件。我的主要问题是当我使用 ctrl + f 并输入我的正则表达式到搜索栏时,它匹配我想要匹配的字符串,但在代码中,它工作错误。这是我的代码:

import re

pattern = re.compile('(([A-Z])[a-z\s,]*)((: ["‘][a-z,!?\.\s]*["’][.,!?])|(; [a-zA-Z\s]*[!.?])|(\s["‘][a-z,.;!?\s]*["’])|([\.?!]))')
text=open('validSentences',"w+")
with open('sentences.txt',encoding='utf8') as file:
    lines = file.readlines()
    for line in lines:
        matches = pattern.fullmatch(line)
        if(matches==None):
            text.write("not valid"+"\n")
        else:
            text.write("valid"+"\n") 
    file.close()

在文档中,它说 fullmatch 仅匹配整个字符串匹配,这就是我正在尝试做的事情,但这段代码对我拥有的所有句子都无效。
我拥有的文本文件:

How can you say that to me? 
As he looked at his reflection in the mirror, he took a deep breath. 
He nodded at himself and, feeling braver, he stepped outside the bathroom. He bumped straight into the 
extremely tall man, who was waiting by the door. 
David said ‘Oh, sorry!’. 
The happy pair discussed their future life 2gether and shared sweet words of admiration. 
We will not stop you; I promise! 
Come here ASAP! 
He pushed his chair back and went to the kitchen at 2 pM. 
I do not know... 
The main character in the movie said: "Play hard. Work harder." 

当我使用 ctrl+f 在 vs 代码中输入我的正则表达式时,整个第一、第二、第四、第七和八行都是高亮的,因此根据fullmatch() 功能,他们需要打印为“有效”,但事实并非如此。我需要有关此问题的帮助。

【问题讨论】:

  • 您的正则表达式可能缺少行尾
  • 首先,删除lines = file.readlines()。然后,for line in lines: 保持尾随换行符,所以使用line=line.rstrip() 或类似的东西。或者确保您的模式以\n 或\n? 甚至\s* 结尾。你会得到 3 个句子,如here。
  • 感谢您的帮助,现在可以使用了
  • 我还有一个问题。我用大写字母开始我的正则表达式,但它匹配以小写字母开头的句子。我不明白为什么

标签: python regex


【解决方案1】:

首先,删除lines = file.readlines(),因为它已经将文件句柄移动到文件流的末尾。然后,您需要记住,在使用for line in lines: 时,line 变量有一个尾随换行符,所以

  • 在运行正则表达式之前使用line=line.rstrip() 删除尾随空格或
  • 确保您的模式以 \n?(可选换行符)或什至 \s*(任何零个或多个空格)结尾。

所以,一个可能的解决方案看起来像

with open('sentences.txt',encoding='utf8') as file:
    for line in file:
        matches = pattern.fullmatch(line.rstrip('\n'))
...

或者,

pattern = re.compile(r'([A-Z][a-z\s,]*)(?:: ["‘][a-z,!?\.\s]*["’][.,!?]|; [a-zA-Z\s]*[!.?]|\s["‘][a-z,.;!?\s]*["’]|[.?!])\s*')
#...
with open('sentences.txt',encoding='utf8') as file:
    for line in file:
....

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-09-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-02-27
    • 1970-01-01
    相关资源
    最近更新 更多