【问题标题】:Parse string based on recurring substring python基于重复出现的子串python解析字符串
【发布时间】:2017-02-07 22:54:35
【问题描述】:

为简洁起见,输入是一个长子字符串:

'这只是虚拟头信息 \n 有时更长,有时更短 \n 日期:1/1/2000 \n 时间:16:00:30 \n 测试名称:meow \n 周期:1 \n 10 15 3 \n 3 69 23 \n 233 33.440 2 \n 频道:HBO \n 周期:1 \n 3 4 5 3 \n 2 3 4 5'

*注意,浮动是制表符分隔的。

我想根据“循环:”进行解析。有时字符串出现一次,有时出现三次。重要数据总是在 Cycle 之后并以一个空白的新行结束。结果可能是列出的数据列表,例如:

[[10 15 3 3 60 23 233 33.440 2], [3 4 5 3 2 3 4 5]]

提前致谢。

编辑

stringlist = re.split(r'\t+', rawsqlstring)
long_data = []
cycle_regex = re.compile('Cycle: ')

for el in stringlist:
    if channel_regex.match(el):
        break
for el in stringlist:
    if el.strip() == '\n':
        break
    long_data.append(el)

【问题讨论】:

  • 你知道正则表达式吗?
  • 是的,我用过几次。埃里克,你的方法是什么?

标签: python regex parsing substring


【解决方案1】:

如果input 是一个字符串

# Split string by newline character, filter out the ones which don't start with 'Cycle: '
# and cut off the 'Cycle' part
cycles = [sub[7:] for sub in input.split('\n') if sub.startswith('Cycle: ')]

# Parse those substrings into lists of ints
nums = [[int(n) for n in c.split(' ')] for c in cycles]

应该有效

您发布的数据很难阅读,但使用 split 的终止字符应该会产生结果

【讨论】:

    猜你喜欢
    • 2015-02-26
    • 1970-01-01
    • 2017-02-15
    • 1970-01-01
    • 1970-01-01
    • 2019-08-26
    • 2011-10-22
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多