【问题标题】:Combine pairs of matching items to matching groups (3+ items)将成对的匹配项目组合到匹配组(3 个以上项目)
【发布时间】:2021-08-05 06:23:06
【问题描述】:

所以,我有一个带有成对匹配图像(或多或少相同)的 csv 文件,如下所示:

File1;File2
File3;File4
File5;File1
File2;File5
...

我想根据这些数据删除重复的图像。 csv 文件来自第三方程序,该程序根据内容匹配图像,但问题是它只匹配对,如果文件之前匹配过,则忽略(即三个以上重复的组没有一起列出)

关键点是上例中的第三行和第四行,其中File1也匹配File5,File2匹配File5,所以三个文件都是重复的。但如果我只是根据数据删除文件,我可能会删除所有文件(如果我删除每行的第二个)。

所以我正在尝试列出一组匹配文件,以确保我始终保留其中一个。上面的例子应该是这样的:

File1;File2;File5
File3;File4
...

这是我的代码,似乎没有这样做。 csv 文件被读入成对列表 (_data)。最后的列表是 self.__map。

    _exists = [ ]
    for _match in _data:
        _first = _match[ 0 ] in _exists
        _second = _match[ 1 ] in _exists
        if not _first and not _second:
            self.__map.append( _match )
            _exists.append( _match[ 0 ] )
            _exists.append( _match[ 1 ] )
        elif _first and not _second:
            for _node in self.__map:
                if _first in _node:
                    _node.append( _match[ 1 ] )
                    _exists.append( _match[ 1 ] )
                    break
        elif _second and not _first:
            for _node in self.__map:
                if _second in _node:
                    _node.append( _match[ 0 ] )
                    _exists.append( _match[ 0 ] )
                    break

看不出为什么它不起作用,但是当我检查 self.__map 时,它不包含 csv 文件中的所有文件。

可能有一种更简单的方法可以做到这一点,所以请随意提出更好的方法。

【问题讨论】:

  • 您可能想使用像networkx 这样的库来查找循环。
  • 谢谢,我怀疑会有类似networkx的东西存在。我对它不够精通,无法掌握我需要的东西,但也许我设法在stackoverflow.com/a/6206011/4038380 找到使用它解决的相同类型的问题

标签: python python-3.x list duplicates combinations


【解决方案1】:

由于这两行,代码不起作用。

if _first in _node:

和

if _second in _node:

一开始_first 和_second 被评估为布尔值,表示第一个文件和第二个文件是否已经被识别。但是在这些行中,它们分别被视为第一项和第二项。将_first 更改为_match[0],将_second 更改为_match[1]。它应该可以工作。

否则我们也可以这样实现。

files = [
    ["File1","File2"],
    ["File3","File4"],
    ["File5","File1"],
    ["File2","File5"],
]

# list of sets. each set contains duplicate/related images. set is used to avoid duplicates. we
# can use list of list here as well. in that case we need add additional logic to handle
# duplicates later
map_of_duplicates = []
for file1, file2 in files:
    new_entry = True
    for match in map_of_duplicates:
        if file1 in match and file2 in match:
            # if both files are present in the match set do nothing
            new_entry = False
            break
        elif file1 in match:
            # if only file1 is present in the match set then add file2 to match set
            new_entry = False
            match.add(file2)
            break
        elif file2 in match:
            #if only file2 is present in the match set then add file1 to the match set
            new_entry = False
            match.add(file1)
            break
    # if after the inner for loop new_entry flag is still true that means none of the files were found in the map and
    # this is indeed a new entry
    if new_entry:
        map_of_duplicates.append(set((file1, file2)))

print(map_of_duplicates)

【讨论】:

    【解决方案2】:

    您可以尝试以下方法:

    data = """File1;File2
    File3;File4
    File5;File1
    File2;File5"""
    
    # Make sure you have a list of set (and not a list of list)
    l = [set(_.split(";")) for _ in data.split("\n")]
    print(l)
    # [{'File1', 'File2'}, {'File4', 'File3'}, {'File1', 'File5'}, {'File5', 'File2'}]
    
    # Add first item
    out = [l[0]]
    # Iterate over the list input
    for pair in l[1:]:
        added = False
        # Iterate over output (already visited pairs)
        for i, pair_out in enumerate(out):
            # if any element of the pair has already been visited
            if any(file_ in pair_out for file_ in pair):
                # Update output
                out[i] = pair_out.union(pair)
                added = True
                # Exit loop
                break
        # If pair not found: append the pair
        if not added: out.append(pair)
    
    print(out)
    # [{'File1', 'File5', 'File2'}, {'File4', 'File3'}]
    

    【讨论】:

      【解决方案3】:

      另一种方法

      from collections import defaultdict
      
      input_data = """File1;File2
      File3;File4
      File5;File1
      File2;File5"""
      
      data = [d.split(';') for d in input_data.split("\n")]
      result = defaultdict(list)
      
      # Create the full list of matches
      for d1, d2 in data:
          result[d1].append(d2)
          result[d2].append(d1)
      
      # Start with all the keys
      keep = list(result.keys())
      
      # Remove any keys that are repeats
      for key in result:
          if key in keep:
              for repeated in result[key]:
                  if repeated in keep:
                      keep.remove(repeated)
      
      # Pick only the ones we want to keep
      result = {k: v for k, v in result.items() if k in keep}
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2015-08-04
        • 2016-07-23
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2015-08-09
        相关资源
        最近更新 更多