【问题标题】:Regular expression to match anything between two symbols正则表达式匹配两个符号之间的任何内容
【发布时间】:2016-10-22 01:45:52
【问题描述】:

试图更多地了解 Python 中的正则表达式,我发现很难匹配两个符号之间的任何字符(包括换行符、制表符、空格等),包括那些符号。

例如:

  • foobar89\n\nfoo\tbar; '''blah blah blah'8&^"'''需要匹配''blah blah blah'8&^"'''

  • fjfdaslfdj; '''blah\n blah\n\t\t blah\n'8&^"'''需要匹配'''blah\n blah\n\t\t blah\n'8&^"'''

(注意,\n 和 \t 符号表示文本文件中的换行符和制表符)

按照这个question,我试过这个^.*\'''(.*)\'''.*$和这个*?\'''(.*)\'''.*,但没有成功。

有人可以指导我做错了什么吗?我也将不胜感激。

另外,为了理解特殊字符转义的概念,我想知道我是否通过替换两个符号(例如从'''到"""或到***)它仍然可以工作的正则表达式(对于相关字符串)?

例如对于

  • fjfdaslfdj; """blah\n blah\n\t\t blah\n'8&^"""需要匹配"""blah\n blah\n\t\t blah\n'8&^"""

更新

我正在尝试测试正则表达式的代码(取自 here 并修改):

import collections
import re

Token = collections.namedtuple('Token', ['typ', 'value', 'line', 'column'])

def tokenize(code):
    token_specification = [
        # regexes suggested from [Thomas Ayoub][3]
        ('BOTH',      r'([\'"]{3}).*?\2'), # for both triple-single quotes and triple-double quotes
        ('SINGLE',    r"('''.*?''')"),     # triple-single quotes 
        ('DOUBLE',    r'(""".*?""")'),     # triple-double quotes 
        # regexes which match OK
        ('COM',       r'#.*'),
        ('NUMBER',  r'\d+(\.\d*)?'),  # Integer or decimal number
        ('ASSIGN',  r':='),           # Assignment operator
        ('END',     r';'),            # Statement terminator
        ('ID',      r'[A-Za-z]+'),    # Identifiers
        ('OP',      r'[+\-*/]'),      # Arithmetic operators
        ('NEWLINE', r'\n'),           # Line endings
        ('SKIP',    r'[ \t]+'),       # Skip over spaces and tabs
        ('MISMATCH',r'.'),            # Any other character
    ]

    test_regexes = ['COM', 'BOTH', 'SINGLE', 'DOUBLE']

    tok_regex = '|'.join('(?P<%s>%s)' % pair for pair in token_specification)
    line_num = 1
    line_start = 0
    for mo in re.finditer(tok_regex, code):
        kind = mo.lastgroup
        value = mo.group(kind)
        if kind == 'NEWLINE':
            line_start = mo.end()
            line_num += 1
        elif kind == 'SKIP':
            pass
        elif kind == 'MISMATCH':
            pass
        else:
            if kind in test_regexes:
                print(kind, value)
            column = mo.start() - line_start
            yield Token(kind, value, line_num, column)

f = r'C:\path_to_python_file_with_above_examples'

with open(f) as sfile:
    content = sfile.read()

for t in tokenize(content):
    pass #print(t)

【问题讨论】:

    标签: regex python-3.x


    【解决方案1】:

    你可以选择:

    ((['"]{3}).*?\2)
    

    见live running python或live running regex


    • ^.*\'''(.*)\'''.*$ => 您在行首/行尾添加了锚点,这在需要多行匹配的情况下不起作用
    • *?\'''(.*)\'''.* => 语法错误
    • re.compile(ur'(([\'"]{3}).*?\2)', re.MULTILINE | re.DOTALL) => re.DOTALL 使 . 匹配新行。

    【讨论】:

    • 它确实适用于''',但不适用于""",我不明白为什么。您能否简要解释一下为什么我尝试的 2 个正则表达式不起作用(特别是因为我从另一个他们正在工作的答案中得到它们)?另外,为什么我不需要“转义”符号(即\''')?
    • 感谢您的编辑和解释。您能否还简要介绍一下其他两件事(匹配其他符号,如""" 和字符转义)? ..我刚刚编辑了我上面的评论和我的问题,为"""添加了一个类似的输入输出示例@
    • @nk-fford 您需要转义作为字符串分隔符的特殊字符:"string with \" as delimiter don't need to escape ' char" 和 'string with \' as delimiter don\'t need to escape " char"。因此,请参阅正则表达式中的转义。如果您需要捕获其他符号,您可以使用character class
    • 人物类很有意思。只是为了确保我做对了,假设我想匹配来自 "dasd a'\n\t """bl"ah"\n 'blah'\n\t\t'8&amp;^""" dasd a 的 """bl"ah"\n 'blah'\n\t\t'8&amp;^""",你能显示一个正则表达式吗?
    • @nk-fford 试试here。如果成功,请转到 代码生成器 并查看 python。如果你失败了,回到我身边;)
    猜你喜欢
    • 2022-07-06
    • 2018-04-04
    • 1970-01-01
    • 2016-11-03
    • 2011-09-06
    • 2018-09-02
    • 1970-01-01
    • 1970-01-01
    • 2011-10-13
    相关资源
    最近更新 更多