【问题标题】:Counting co-occurence in a list of list计算列表列表中的共现
【发布时间】:2021-07-12 22:54:34
【问题描述】:

使用python,我定义了以下字符串列表

some_list = [['abc', 'aa', 'xdf'], ['def', 'asd'], ['abc', 'xyz'], ['ghi', 'edd'], ['abc', 'xyz'], ['abc', ]]

选择和计算出现在与 'abc' 相同的子列表中的字符串的最佳方法是什么?在这个例子中,我正在寻找一个输出(字典或列表),例如:

('aa', 1), ('xdf', 1), ('xyz',2), (None, 1)

(None, 1) 将捕获 'abc' 是子列表的唯一字符串的情况。

【问题讨论】:

  • 当None 没有出现在任何列表中时,您希望如何在结果列表中返回(None, 1)?当'abc' 被隔离在它自己的列表中时会出现这种情况吗?
  • 是的,我们的想法是同时捕获“abc”作为子列表的单个元素的次数。 (抱歉,None 不清楚)

标签: python list counter


【解决方案1】:

我会先把它们都数一遍,然后分别处理边缘情况:

from collections import Counter
from itertools import chain

c = Counter(chain.from_iterable(x for x in A if 'abc' in x))
c[None] = A.count(['abc'])

【讨论】:

  • 对于短子列表的长列表,推荐哪种解决方案,使用Counter和itertools,还是@rahlf23提出的解决方案(使用for循环)?
  • @djourd1:除非您需要每秒检查数百万个元素数百次,否则这并不重要。
【解决方案2】:

这是一种利用defaultdict() 的长期方法:

from collections import defaultdict

some_list = [['abc', 'aa', 'xdf'], ['def', 'asd'], ['abc', 'xyz'], ['ghi', 'edd'], ['abc', 'xyz'], ['abc', ]]

d = defaultdict(int)

for i in some_list:
    if 'abc' in i:
        if len(i)==1:
            d[None] += 1
        for j in i:
            if j=='abc': continue
            d[j] += 1

产量:

defaultdict(<class 'int'>, {'aa': 1, 'xdf': 1, 'xyz': 2, None: 1})

【讨论】:

    【解决方案3】:

    您可以根据搜索条件过滤您的列表,然后过滤掉搜索值以获得您的同现。然后你可以使用任何方法来计算它们的频率。

    from collections import Counter
    from itertools import chain
    
    some_list = [['abc', 'aa', 'xdf'], ['def', 'asd'], ['abc', 'xyz'], ['ghi', 'edd'], ['abc', 'xyz'], ['abc', ]]
    search = 'abc'
    co_occurrences = [[v for v in lst if v != search] if len(lst) > 1 else [None] for lst in some_list if search in lst]
    print(co_occurrences)
    c = Counter(chain.from_iterable(co_occurrences))
    print(c)
    

    输出:

    [['aa', 'xdf'], ['xyz'], ['xyz'], [None]]
    Counter({'xyz': 2, 'aa': 1, 'xdf': 1, None: 1})
    

    【讨论】:

      【解决方案4】:

      我的方法没有导入任何包:

      some_list = [['abc', 'aa', 'xdf'], ['def', 'asd'], ['abc', 'xyz'], 
                   ['ghi', 'edd'], ['abc', 'xyz'], ['abc', ]]
      d = {'None':0}
      for e in some_list:
          if 'abc' in e and len(e) != 1:
              for f in e:
                  if f != 'abc' and f not in d:
                      d[f] = 1
                  elif f != 'abc' and f in d:
                      d[f] +=1
          elif 'abc' in e and len(e) == 1:
              d['None'] += 1
      print(d)
      

      此代码将打印:

      {'None': 1, 'aa': 1, 'xdf': 1, 'xyz': 2}
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-07-16
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多