【问题标题】:Splitting a CSV file into multiple csv by target columns values按目标列值将 CSV 文件拆分为多个 csv
【发布时间】:2019-02-14 20:12:13
【问题描述】:

总的来说,我对编程和 Python 还很陌生。我有一个很大的 CSV 文件,我需要根据目标列(最后一列)的目标值将其拆分为多个 CSV 文件。

这是我要拆分的 CSV 文件数据的简化版本。

1254.00   1364.00   4562.33   4595.32   1
1235.45   1765.22   4563.45   4862.54   1
6235.23   4563.00   7832.31   5320.36   1
8623.75   5632.09   4586.25   9361.86   0
5659.92   5278.21   8632.02   4567.92   0
4965.25   1983.78   4326.50   7901.10   1
7453.12   4993.20   4573.30   8632.08   1
8963.51   7496.56   4219.36   7456.46   1
9632.23   7591.63   8612.37   4591.00   1
7632.08   4563.85   4632.09   6321.27   0
4693.12   7621.93   5201.37   7693.48   0
6351.96   7216.35   795.52    4109.05   0

我想拆分,以便输出提取不同 csv 文件中的数据,如下所示:

sample1.csv

1254.00   1364.00   4562.33   4595.32   1
1235.45   1765.22   4563.45   4862.54   1
6235.23   4563.00   7832.31   5320.36   1

sample2.csv

8623.75   5632.09   4586.25   9361.86   0
5659.92   5278.21   8632.02   4567.92   0

sample3.csv

4965.25   1983.78   4326.50   7901.10   1
7453.12   4993.20   4573.30   8632.08   1
8963.51   7496.56   4219.36   7456.46   1
9632.23   7591.63   8612.37   4591.00   1

sample4.csv

7632.08   4563.85   4632.09   6321.27   0
4693.12   7621.93   5201.37   7693.48   0
6351.96   7216.35   795.52    4109.05   0

我尝试使用 pandas 和一些 groupby 函数,但它将所有 1 和 0 合并到单独的文件中,其中一个包含所有值为 1 和另一个 0 的值,这不是我需要的输出。

任何帮助将不胜感激。

【问题讨论】:

  • 你试过什么?每次当最后一列中的值发生变化时,只需遍历文件并开始写入新文件...

标签: python csv


【解决方案1】:

您可以做的是获取每行中最后一列的值。如果该值与前一行中的值相同,则将该行添加到同一个列表中,如果不只是创建一个新列表并将该行添加到该空列表中。对于数据结构,使用列表列表。

【讨论】:

    【解决方案2】:

    假设文件 'input.csv' 包含原始数据。

    1254.00   1364.00   4562.33   4595.32   1
    1235.45   1765.22   4563.45   4862.54   1
    6235.23   4563.00   7832.31   5320.36   1
    8623.75   5632.09   4586.25   9361.86   0
    5659.92   5278.21   8632.02   4567.92   0
    4965.25   1983.78   4326.50   7901.10   1
    7453.12   4993.20   4573.30   8632.08   1
    8963.51   7496.56   4219.36   7456.46   1
    9632.23   7591.63   8612.37   4591.00   1
    7632.08   4563.85   4632.09   6321.27   0
    4693.12   7621.93   5201.37   7693.48   0
    6351.96   7216.35   795.52    4109.05   0
    

    代码如下

    target = None
    counter = 0
    with open('input.csv', 'r') as file_in:
        lines = file_in.readlines()
        tmp = []
        for idx, line in enumerate(lines):
            _target = line.split(' ')[-1].strip()
            if idx == 0:
                tmp.append(line)
                target = _target
                continue
            else:
                last_line = idx + 1 == len(lines)
                if _target != target or last_line:
                    if last_line:
                        tmp.append(line)
                    counter += 1
                    with open('sample{}.csv'.format(counter), 'w') as file_out:
                        file_out.writelines(tmp)
                    tmp = [line]
                else:
                    tmp.append(line)
                target = _target
    

    【讨论】:

    • 您的输出与我的错误输出相似。我不希望我的输出只有 2 个文件,一个包含 0,另一个包含 1。您可以检查我在问题中给出的所需输出。还是谢谢!
    • 您已要求根据最后一列的“目标”创建输出文件。这就是代码的作用。请说明当您提供的数据源中只有两个目标 [0,1] 时,应该如何使代码创建像 sample3.csv 等文件。
    • 感谢您的努力。我想要的迭代是,在数据中,我们看到前 3 行的目标值为 1。因此,sample1.csv 文件应该包含前 3 行。当目标值从 1 变为 0 时,它应该创建一个新的 sample2.csv,其中接下来的两行包含目标值 0。然后当迭代器发现目标值从 0 变为 1 时(在第 6 行),它应该创建一个新的 sample3.csv 并将下一行的目标值设置为 1,依此类推。我希望它澄清。请看看我原来的问题。我已经解释了我想要的输出。谢谢!
    • 好的...明白了。代码已修改。看看吧。
    • 谢谢。我试过你修改过的。但我得到 12 个 sample.csv 文件而不是 4 个 csv 文件。因为现在每行都被创建为 csv 文件,而不是目标值组。 Your output: sample1.csv 1254.00 1364.00 4562.33 4595.32 1 sample2.csv 1235.45 1765.22 4563.45 4862.54 1 Where I wanted: sample1.csv 1254.00 1364.00 4562.33 4595.32 1 1235.45 1765.22 4563.45 4862.54 1 6235.23 4563.00 7832.31 5320.36 1
    【解决方案3】:

    也许你想要这样的东西:

    from itertools import groupby
    from operator import itemgetter
    
    sep = '   '
    
    with open('data.csv') as f:
        data = f.read()
    
    split_data = [row.split(sep) for row in data.split('\n')]
    gb = groupby(split_data, key=itemgetter(4))
    
    for index, (key, group) in enumerate(gb):
        with open('sample{}.csv'.format(index), 'w') as f:
            write_data = '\n'.join(sep.join(cell) for cell in group)
            f.write(write_data)
    

    与pd.groupby 不同,itertools.groupby 不会预先对源进行排序。这会将输入 CSV 解析为列表列表,并根据包含目标的第 5 列对外部列表执行 groupby。 groupby 对象是组的迭代器;通过将每个组写入不同的文件,可以达到您想要的结果。

    【讨论】:

    • 谢谢。但我收到了第 10 行的索引错误。IndexError: list index out of range
    • @MishkatRahman 问题可能是我给出的代码假定源文件的格式与您在问题中所说的完全相同(元素之间有 3 个空格。)如果它真的是您所说的 CSV ,您需要将 sep 值更改为其他值。
    • 好的。感谢马库斯的所有帮助!
    【解决方案4】:

    我建议使用一个函数来做被要求的事情。

    可能会留下未引用的文件对象 我们已经打开了写作,所以当它们自动关闭时 垃圾收集,但在这里我更喜欢明确关闭每个输出 文件,然后再打开另一个文件。

    该脚本被大量评论,因此没有进一步的解释:

    def split_data(data_fname, key_len=1, basename='file%03d.txt')
    
        data = open(data_fname)
    
        current_output = None # because we have yet not opened an output file
        prev_key = int(1)     # because a string is always different from an int
        count = 0             # because we want to count the output files
    
        for line in data:
    
            # line has a trailing newline so that to extract the key
            # we have to take into account that
            key = line[-key_len-1:-1]
    
            if key !=  prev_key     # key has changed!
    
               count += 1           # a new file is going to be opened
               prev_key = key       # remember the new key
               if current_output:   # if a file was opened, close it
                   current_output.close()
               # open a new output file, its name derived from the variable count
               current_output = open(basename%count, 'w')
    
            # now we can write to the output file
            current_output.write(line)
            # note that line is already newline terminated
    
        # clean up what is still going
        current_output.close()
    

    这个答案有an history。

    【讨论】:

    • 感谢 gboffi 的所有解释。请问,在您的修改版本中,我应该如何处理 f.write(line)?因为我们没有任何参考 f。在您之前的代码中,f = None。
    • 我在重构时忘记了名称转换...哎呀! — 当然它应该是current_output.write(line),因为这是我们要写我们正在处理的行的地方。我已经编辑了答案,对错误和随之而来的困惑表示歉意。
    猜你喜欢
    • 1970-01-01
    • 2017-01-27
    • 1970-01-01
    • 2018-06-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-08-21
    相关资源
    最近更新 更多