【问题标题】:Detect and extract changing numbers in list of strings检测并提取字符串列表中的变化数字
【发布时间】:2013-06-06 10:43:49
【问题描述】:

假设我有一个音频文件名列表(它可以是任何带有连续数字的字符串列表),它们具有不同的命名方案,但所有文件名中都包含轨道号。

我想提取变化的数字。

示例 1

Fooband 41 - Live - 1. Foo Title
...
Fooband 41 - Live - 11. Another Foo Title

想要的结果

号码列表:1,2,3,...,11

示例 2

02. Barband - Foo Title with a 4 in it
05. Barband - Another Foo Title
03. Barband - Bar Title
...
17. Barband - Yet another Foo Title

想要的结果

号码列表:2,5,3,...,17

由于索引号的位置不固定,我(认为)我不能在那里使用正则表达式。

我有什么

  1. 找到字符串的共同前缀和后缀并将其删除
  2. 查看字符串左右两边是否有数字
  3. 使用该数字获取索引

但有一个问题:如果我找到 示例 1 的通用前缀,那么通用前缀将是 Fooband 41 - Live - 1,所以 1 会丢失(对于像 Song X - 10, Song X - 11, ...) 这样的命名方案也是如此。

问题

有什么好方法可以检测和提取字符串列表中不断变化的数字(在相似位置)?

我正在使用 Python(这对这个问题并不重要)

如果我也能检测到罗马数字,那将是一个好处,但我怀疑这会更困难。

【问题讨论】:

  • 我认为你应该有一个直接的结构来命名你的音频文件名,这样你就可以得到所有情况的解决方案。
  • 是的,但我只是想说明音频文件的问题以便于理解(在命名音频文件时,我实际上非常迂腐;-)。解决方案应该适用于任何字符串列表。

标签: python algorithm language-agnostic


【解决方案1】:
f = open('data.txt')
data = []

pattern = "\d+|[IVX]+"
regex = re.compile(pattern)

for line in f:
    matches = re.findall(regex, line)
    data.append(matches)

f.close()

print data
transposed_data = zip(*data)
print transposed_data

for atuple in transposed_data:
    val = atuple[0]

    if all([num==val for num in atuple]): 
        next
    else:
        print atuple
        break

数据.txt:

Fooband 41 - Live - 1. Foo Title
Fooband 41 - Live - 2. Foo Title
Fooband 41 - Live - 3. Foo Title
Fooband 41 - Live - 11. Another Foo Title

--输出:--

[['41', '1'], ['41', '2'], ['41', '3'], ['41', '11']]
[('41', '41', '41', '41'), ('1', '2', '3', '11')]
('1', '2', '3', '11')

数据.txt:

01. Barband - Foo Title with a 4 in it
05. Barband - Another Foo Title
03. Barband - Bar Title
17. Barband - Yet another Foo Title

--输出:--

[['01', '4'], ['05'], ['03'], ['17']]
[('01', '05', '03', '17')]
('01', '05', '03', '17')

数据.txt:

01 Barband - Foo Title with a (I) in it
01 Barband - Another Foo (II) Title
01. Barband - Bar Title (IV)
01. Barband - Yet another (XII) Foo Title

--输出:--

[['01', 'I'], ['01', 'II'], ['01', 'IV'], ['01', 'XII']]
[('01', '01', '01', '01'), ('I', 'II', 'IV', 'XII')]
('I', 'II', 'IV', 'XII')

【讨论】:

  • 花了我一两秒钟才弄明白,但这很好!
【解决方案2】:

如果它们的格式相似,您可以使用 python 的re module。从字符串列表中提取这些数字的简短代码如下所示:

import re
regex = re.compile(".*([0-9]+).*")

number = regex.match("Fooband 41 - Live - 1. Foo Title").group(1)

【讨论】:

    猜你喜欢
    • 2011-08-08
    • 1970-01-01
    • 2014-11-09
    • 2023-03-19
    • 1970-01-01
    • 1970-01-01
    • 2016-12-17
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多