【问题标题】:How to capture all content between two captured groups如何捕获两个捕获组之间的所有内容
【发布时间】:2017-10-11 03:18:25
【问题描述】:

我有一个 txt 文件,它是从包含一长串项目的 pdf 转换而来的。这些项目的编号约定如下:

[A-Z]{1,2}\d{1,2}\.\d{1,2}\.\d{1,2}

此表达式将匹配以下内容:

A1.1.1

和

ZZ99.99.99

这很好用。我遇到的问题是我试图在第 1 组中捕获此内容以及第 2 组中每个项目编号(项目描述)之间的所有内容。

我还需要将这些作为列表或迭代返回,以便最终将捕获的内容导出到 Excel 电子表格。

这是我目前的正则表达式:

^([A-Z]{1,2}\d{1,2}\.\d{1,2}\.\d{1,2}\s)([\w\W]*?)(?:\n)

点击此链接以查找我所拥有的和面临的问题的示例:

Debuggex Demo

有没有人能帮我弄清楚如何捕捉每个数字之间的所有内容,无论有多少段?

任何意见将不胜感激,谢谢!

【问题讨论】:

  • 我不懂 Python,但我最近有一个类似的question。这是regex101 demo。希望对你有帮助

标签: python regex


【解决方案1】:

你很亲密:

import re

s = """
A1.2.1 This is the first paragraph of the description that is being captured by the regex even if the description contains multiple lines of text.ZZ99.99.99
"""
final_data = re.findall("[A-Z]{1,2}\d{1,2}\.\d{1,2}\.\d{1,2}(.*?)[A-Z]{1,2}\d{1,2}\.\d{1,2}\.\d{1,2}", s)

输出:

[' This is the first paragraph of the description that is being captured by the regex even if the description contains multiple lines of text.']

通过使用(.*?),您可以匹配第一个正则表达式定义的字母和数字之间的任何文本。

【讨论】:

猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2012-11-06
  • 2019-01-16
  • 1970-01-01
  • 1970-01-01
  • 2022-06-15
  • 2021-11-11
  • 2017-06-11
相关资源
最近更新 更多