【问题标题】:Extract all substrings between two markers for a very long string提取两个标记之间的所有子字符串以获得非常长的字符串
【发布时间】:2020-06-12 11:18:31
【问题描述】:

这是问题Extract all substrings between two markers 的延续。 @Daweo 和@Tim Biegeleisen 的answers 适用于小弦乐。

但是对于非常大的字符串,正则表达式似乎不起作用。这可能是由于字符串长度的限制,如下所示:

>>> import re
>>> teststr = "&marker1\nThe String that I want /\n&marker1\nAnother string that I want /\n"
>>> for i in range(0, 23):
...    teststr += teststr # creating a very long string here
... 
>>> len(teststr)
603979776
>>> found = re.findall(r"\&marker1\n(.*?)/\n", newstr)
>>> len(found)
46
>>> found
['The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ', 'The String that I want ', 'Another string that I want ']

我可以做些什么来解决这个问题并找到制造商 start="&maker1" 和 end="/\n" 之间的所有事件? re可以处理的最大字符串长度是多少?

【问题讨论】:

  • 可以在我的机器上运行,至少当我将newstr 替换为teststr 时。
  • @Ronald 该问题已被编辑。还能用吗?
  • 它适用于我的家用机器,使用 Python 3.8.3。 len(found) 打印为16777216。
  • 仍然有效 ;-) 我想把它推到极限取决于你系统的内存?但是你得到的错误是什么?

标签: python python-3.x python-2.7 python-re


【解决方案1】:

我无法让re.findall 工作。现在我确实使用re,但要查找标记的位置并手动提取子字符串。

locs_start = [match.start() for match in re.finditer("\&marker1", mylongstring)]
locs_end = [match.start() for match in re.finditer("/\n", mylongstring)]

substrings = []
for i in range(0, len(locs_start)):
    substrings.append(mylongstring[locs_start[i]:locs_end[i]+1])

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-06-22
    • 2011-06-07
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多