【问题标题】:Python - Compare first item in sublist, if repeated, compare third item and pick lesser valuePython - 比较子列表中的第一项,如果重复,则比较第三项并选择较小的值
【发布时间】:2015-08-15 13:54:22
【问题描述】:

所以我有一个包含子列表的列表。

例如

biglist = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Hello', 'Friend', '1.5'], 
['Blue', 'Yo', 'Dude', '1.2'], ['Green', 'Bon', 'Jour', '0.1'], 
['Purple', 'Hey', 'Sup', '0.4'], ['Purple', 'Greetings', 'Pal', '2.8']]

这就是我想要做的......我想通过这个迭代来执行以下操作:

  • 对于每个子列表,读取位置 0。
  • 如果在位置 0 存在另一个具有相同字符串的子列表,则读取位置 3
  • 无论哪个数字在位置 3 中较低,将另一个子列表完全删除并保留具有较小值的子列表。有时有两个以上的子列表具有相同的位置[0]

所以,对于我的示例列表。我想保留“红色”子列表,比较两个“蓝色”子列表并将数值较小的子列表保留在特定位置 3,然后还保留“绿色”子列表。我一直在搞乱 set() 但有点难过。起初我尝试将它散列到 0 位置(红色、蓝色等)是关键的位置,其余位置是值(作为列表),但我被卡住了,走了一条不同的路线。

想要的结果:

biglist = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Yo', 'Dude', '1.2'], 
['Green', 'Bon', 'Jour', '0.1'], ['Purple', 'Hey', 'Sup', '0.4']]

注意:我正在使用的列表是由前一个函数传递的。

我在另一个问题上找到了这个,但是 set() 让我有点困惑,我不知道如何进一步研究第三个位置或如何正确传递我已经通过另一个函数创建的列表在这个之前在同一个脚本中。当我在尝试传递列表时运行它时,我什么也得不到。

def unique_items(L):
found = set()
for item in L:
    if item[0] not in found:
        yield item
        found.add(item[0])

非常感谢。

【问题讨论】:

  • 我认为在原始列表和期望列表中,'Hi' 'There' 之间应该有逗号,对吗?请相应地进行编辑。我不允许编辑,因为更改的数量只有 '2' 而不是编辑所需的 '6' :-)
  • 是的,谢谢,我以为我编辑了它,但一定只是在我的脚本中这样做。

标签: python list set compare sublist


【解决方案1】:

首先,我将创建一个列表字典,其中子列表中的第一项(颜色)作为键,值将是元组(列表中子列表的索引,子列表中的最后一项):

from collections import defaultdict
x = defaultdict(list)

# This for loop extracts the index of each sublist (i) and then
# assigns the contents of the sublist to variables, in this case
# we want the first item in the sublist to be the 'key', ignore
# everything in between and grab the last item as the 'val'.
# If the sublists have arbitrary number of items then you could
# use for i, item in enumerate(biglist) and replace key with
# item[0] and val with item[3]
for i, (key, *_, val) in enumerate(biglist):
    x[key].append((i, float(val))

x 现在看起来像:

defaultdict(<class 'list'>, {'Blue': [(1, '1.5'), (2, '1.2')], 'Purple': [(4, '0.4'), (5, '2.8')], 'Green': [(3, '0.1')], 'Red': [(0, '0.534')]})

然后我会通过

创建一个新列表
  • 按升序对字典 x 中每种颜色的条目进行排序,以便列表中的第一项是具有最小“权重”值的项(您称为位置 3)
  • 将该排序列表的第一项作为包含子列表索引的元组作为其第一项
  • 最终使用索引检索子列表

比如:

res = [biglist[sorted(val, key=lambda x: x[1])[0][0]] for val in x.values()]

res 现在包含

[['Blue', 'Yo', 'Dude', '1.2'],
 ['Purple', 'Hey', 'Sup', '0.4'],
 ['Green', 'Bon', 'Jour', '0.1'],
 ['Red', 'HiThere', '0.534']]

【讨论】:

  • 谢谢我的朋友!我最初有一个类似的字典,但它的值是一个列表,而不是一个元组。你能在这里解释一下for循环吗? “_”让我感到困惑。我得到了无效的语法,所以我认为需要用其他东西代替“_”。 ***实际上,我想我知道为什么这不起作用。数字后面的子列表中实际上有一个项目。例如,从技术上讲,红色实际上是 ['Red', 'Hi' 'There', '0.534', 'Another thing'],所以它不可能是一个元组,因为有多个值。对于我决定保留的子列表,我将需要所有这些项目。对不起,我漏掉了!
  • 好的,所以我使用了 defaultdict 但不同的是得到了这个,现在我会继续更新。 x = defaultdict(list) for biglist 中的项目: x[item[0]].append(item[1:]) print x
  • 我在上面的代码中添加了一些 cmets。如果您的列表包含任意数量的项目,您可以将访问变量的方式更改为您建议的方式。这就是你需要改变的全部。该 dict 只是为了让我们知道我们想要保留哪些子列表。最终结果将包含您要保留的子列表中的所有数据。
【解决方案2】:

这是另一种方法 - 我认为更具可读性

# result_list for verification
result_list = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Yo', 'Dude', '1.2'], ['Green', 'Bon', 'Jour', '0.1'], ['Purple', 'Hey', 'Sup', '0.4']]

# original list 
biglist = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Hello', 'Friend', '1.5'], ['Blue', 'Yo', 'Dude', '1.2'], ['Green', 'Bon', 'Jour', '0.1'],
['Purple', 'Hey', 'Sup', '0.4'], ['Purple', 'Greetings', 'Pal', '2.8']]

another_list = []

import itertools

# Sort the big list by tuple of x[0], x[3] First sort by x[0] and then resolve tie by x[3]
biglist = sorted(biglist, key=lambda x:(x[0],x[3]))

# now group the list by the first element of each list, y gives an iterator, we simply make a list of that and take first element.

for x, y in itertools.groupby(biglist, lambda x:x[0]):
    another_list.append(list(y)[0])

# following line is just for verification
print another_list == sorted(result_list)

注意:此处不保留原始列表中的顺序。如果你想保留它,下面应该工作

 result_list = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Yo', 'Dude', '1.2'],
['Green', 'Bon', 'Jour', '0.1'], ['Purple', 'Hey', 'Sup', '0.4']]
biglist = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Hello', 'Friend', '1.5'],
['Blue', 'Yo', 'Dude', '1.2'], ['Green', 'Bon', 'Jour', '0.1'],
['Purple', 'Hey', 'Sup', '0.4'], ['Purple', 'Greetings', 'Pal', '2.8']]

#print sorted(sorted(biglist), key=lambda x:(x[0],x[3]))
cleanup_list = []
import itertools
s_biglist = sorted(biglist, key=lambda x:(x[0],x[3]))

for x, y in itertools.groupby(s_biglist, lambda x:x[0]):
    cleanup_list.extend(list(y)[1:])

for x in cleanup_list:
    biglist.remove(x)

print biglist

【讨论】:

  • 谢谢!我想出了另一种方式,我会再次发布它,以便人们可以正确看到它
【解决方案3】:

非常感谢大家!这就是我以自己的方式解决它的方法。也许这不是最好的方法,而且我知道我的变量没有被命名为最好的(我实际上在这个例子中故意将它们命名为简单化,但对于我的真实脚本来说,它们是具体和独特的)。实际的输入列表来自一个超过 50,000 行长的文件。

#!/usr/bin/python

from collections import defaultdict

biglist = [['Red', 'Hi', 'There', '0.534'], ['Blue', 'Hello', 'Friend', '1.5'], 
['Blue', 'Yo', 'Dude', '1.2'], ['Green', 'Bon', 'Jour', '0.1'], 
['Purple', 'Hey', 'Sup', '0.4'], ['Purple', 'Greetings', 'Pal', '2.8']]

x = defaultdict(list)
for item in biglist:
    x[item[0]].append(item[1:])

y = dict(x)

for key, value in y.items():
    FINAL = []
    if len(value) <= 1:
        FINAL.append(value)
    else:
        valuelist = []
        for version in value:
            valuelist.append(version[3])
            best = min(valuelist)
        for version in value:
            if version[3] == best:
                FINAL.append(version)
        y[key] = FINAL
print y

【讨论】:

    猜你喜欢
    • 2020-10-24
    • 2022-01-16
    • 2012-09-12
    • 2020-05-17
    • 1970-01-01
    • 2022-01-19
    • 1970-01-01
    • 2020-10-23
    • 2022-01-23
    相关资源
    最近更新 更多