【问题标题】:Merge Columns and Remove Duplication合并列并删除重复
【发布时间】:2015-01-01 19:19:45
【问题描述】:

我有一个包含 2 列数据的输入文件。我需要合并两列并删除重复项。任何建议如何开始?谢谢 !

输入文件

5045 2317
5045 1670
5045 2156
5045 1509
5045 3833
5045 1013
5045 3491
5045 32
5045 1482
5045 2495
5045 4280
5045 1380
5045 3998

预期输出

 5045 
 2317
 1670
 2156
 1509
 3833
 1013
 3491
 32
 1482
 2495
 4280
 1380
 3998

【问题讨论】:

  • @PadraicCunningham 没有顺序无所谓

标签: python file merge duplication


【解决方案1】:
set1 = set()
set2 = set()
for line in myfile:
    a,b = line.strip().split()
    set1.add(int(a))
    set2.add(int(b))
set1.update(set2)

然后将set1的内容写入文件。

【讨论】:

  • 你为什么要排序和投射?
  • 您将失去 set 的订单。你能在 1670 之前授予 2317 打印吗?
  • true,他的预期输出没有排序。我会编辑。
【解决方案2】:

我假设输出中行的顺序确实很重要。下面代码的输出将与您想要的输出完全匹配(例如,与使用sets 的答案不同):

In [1]: with open("file.txt") as f, open("output.txt", "w") as out:
   ...:     arrs = [ l.rstrip().split() for l in f ] 
   ...:     vals = [ a for arr in arrs for a in arr ] # merge columns
   ...:     # restrict to first occurrence of each value (i.e. remove duplicates)
   ...:     uniqueVals = [ v for i, v in enumerate(vals) if vals.index(v) == i ]
   ...:     out.write("\n".join(uniqueVals))

这会从"file.txt" 获取输入并通过以下方式输出到"output.txt":

  1. 正在加载输入文件。
  2. 合并两列。
  3. 限制每个值第一次出现。

【讨论】:

    【解决方案3】:
    >>> import numpy as np
    >>> a=np.loadtxt('file_name',delimiter=' ')
    >>> a=a.flatten()
    >>> a=list(set(a))
    >>> a
    [32.0, 3491.0, 1380.0, 1509.0, 1670.0, 1482.0, 2156.0, 2317.0, 5045.0, 4280.0, 3833.0, 2495.0, 3998.0, 1013.0]
    

    【讨论】:

      【解决方案4】:

      为了保持顺序:

      from itertools import chain
      with open("in.txt") as f:
          lines = list(chain.from_iterable(x.split() for x in f))
          with open("in.txt","w") as f1:
              for ind, line in enumerate(lines,1):
                  if not line in lines[:ind-1]:
                      f1.write(line+"\n")
      

      输出:

      5045
      2317
      1670
      2156
      1509
      3833
      1013
      3491
      32
      1482
      2495
      4280
      1380
      3998
      

      如果顺序无关紧要:

      from itertools import chain
      with open("in.txt") as f:
          lines = set(chain.from_iterable(x.split() for x in f))
          with open("in.txt","w") as f1:
              f1.writelines("\n".join(lines))
      

      如果第一列只有一个数字重复:

      with open("in.txt") as f:
          col_1 = f.next().split()[0] # get first column number
          lines = set(x.split()[1] for x in f) # get all second column nums
          lines.add(col_1) # add first column num
          with open("in.txt","w") as f1:
              f1.writelines("\n".join(lines))
      

      【讨论】:

      • 这很好用..但是当涉及到大文件时,它需要时间
      • 你的文件有多大?
      • 我有 KB、MB 和 GB 的文件范围
      • 我想不出比设置删除副本更有效的方法,第一列总是重复吗?
      • 所以我们只需要第一列的单个值,它总是只包含一个重复的数字?
      猜你喜欢
      • 2016-10-24
      • 1970-01-01
      • 2020-08-16
      • 1970-01-01
      • 2011-11-18
      • 2015-11-29
      • 2017-06-05
      • 1970-01-01
      • 2018-03-11
      相关资源
      最近更新 更多