【问题标题】:write a list to CSV file and start new column if condition is met如果满足条件,将列表写入 CSV 文件并开始新列
【发布时间】:2018-01-18 21:51:20
【问题描述】:

我有一个包含数据点和“标识符”的列表,如下所示:

['identifier', 1, 2, 3, 4, 'identifier', 10, 11, 12, 13, 'identifier', ...]

我想将此列表写入 CSV 文件并为每个标识符创建一个新列。 例如

 for data in list:
        if data=='identifier':
            ==> create a new column in the CSV file and print the subsequent data points

我期待听到您的建议。

干杯,

-塞巴斯蒂安

【问题讨论】:

  • 我期待看到您的尝试。在这个网站上已经有数以千计的关于读/写 CSV 的问题。您从研究中尝试过什么?
  • 嗨,我已经尝试了大部分。我缺少的元素是告诉作者在满足条件时开始新列。
  • 请将您的最佳尝试作为对问题的编辑。如果我们在回答问题时也能解决您的误解,这对您来说会更有用
  • 你可能还会想为什么一个“标识符”后面跟着 4 个值以及为什么它们应该进入新列 - 你从不谈论行...
  • @pault 是的,你可以。 writerows 方法采用嵌套列表,每个内部列表代表一行(该列表中的每个项目都在单独的列中)。您可以使用 for 循环轻松将此输入分解为行和列作为嵌套列表,并可能将其压缩为列表推导。

标签: python csv


【解决方案1】:

此解决方案不会将数据写入 csv 文件,但使用 csv 库这是一个简单的步骤。这样做的作用是将您提供的数据重组为列表列表,每个子列表都是单行数据。

l = ['identifier', 1, 2, 3, 'identifier', 10, 11, 12, 13, 'identifier', 4, 3, 2, 1, 10]

def split_list(l, on):
    """Splits a list an identifier and returns a list of lists split on the
    identifier without including it."""
    splits = []
    cache = []
    for v in l:
        # Check if this is an identifier
        if v == on:
            # Add the cache to splits unless it is empty
            if cache:
                splits.append(cache)
                # Empty the cache
                cache = []
        else:
            cache.append(v)
    # Add the last cache to splits if it is not empyt
    if cache:
        splits.append(cache)
    return splits

def reshape_list(l, default=None):
    """Takes a list of lists assuming each list is a column of values and
    reshapes it to be a list of rows, if list are not all the same length None
    will be used to fill empyt spots."""
    result = []
    # Get the length of the longest list
    maxlen = max(map(len, l))
    for i in range(maxlen):
        # Create each row
        row = []
        # Extract the values from the columns
        for column in l:
            if i < len(column):
                row.append(column[i])
            else:
                row.append(default)
        result.append(row)
    return result


print(l)
t = split_list(l, 'identifier')
print(t)
r = reshape_list(t)
print(r)

【讨论】:

    【解决方案2】:

    生成演示数据:

    import random
    
    random.seed(20180119) # remove to get random data between runs
    id = 'identifier'
    
    def genData():
        data = []
        for n in range(10+random.randint(1,10)):
            data.append(id)
            data.extend(random.choices(range(1,20),k=random.randint(3,12)))
        print(data)
        return data
    

    输出:

    ['identifier', 18, 6, 19, 10, 12, 18, 17, 12, 
     'identifier', 10, 17, 17, 10, 15, 12, 16, 18, 19, 18, 14, 9, 
     'identifier', 6, 10, 1, 14, 4, 
     'identifier', 3, 7, 7, 4, 8, 2, 16, 8, 1, 8, 16, 6, 
     'identifier', 6, 17, 8, 8, 13, 15, 7, 9, 4, 10, 15, 
     'identifier', 17, 8, 3, 8, 2, 19, 16, 2, 5, 6, 
     'identifier', 18, 6, 18, 19, 7, 8, 14, 7, 7, 19, 
     'identifier', 13, 7, 4, 13, 
     'identifier', 15, 8, 17, 8, 1, 12, 16, 7, 5, 19, 14, 9, 
     'identifier', 18, 16, 10, 7, 16, 18, 19, 6, 15, 8, 13, 15, 
     'identifier', 15, 2, 18, 13, 7, 
     'identifier', 17, 19, 15, 4, 18, 7, 13, 17, 8, 9, 
     'identifier', 9, 17, 18, 8, 17, 17, 17, 
     'identifier', 3, 16, 15, 13, 9, 
     'identifier', 15, 12, 2, 16, 2, 5, 16, 18]
    

    重新格式化:

    def partitionData(idToUse,dataToUse):
        lastId = None
        for (i,n) in enumerate(data):       # identify subslices of data
            if n == idToUse and not lastId:     # find first id, data before is discarded
              lastId = i
              continue
    
            if n == idToUse:                    # found id
              yield data[lastId:i]                  # yield sublist including idToUse
              lastId = i
    
        if (data[-1] != id):                    # yield rest of data
            yield data[lastId:]
    

    写入数据:

    data = genData()
    partitioned = partitionData(id, data)
    
    import itertools
    import csv
    with open('result.csv', 'w', newline='') as csvfile:
        writer = csv.writer(csvfile, delimiter=";")
        # like zip, but fills up shorter ones with None till longest index
        writer.writerows(itertools.zip_longest(*partitioned, fillvalue=None)) 
    

    result.csv:

    identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier;identifier
    10;6;3;6;17;18;13;15;18;15;17;9;3;15
    17;10;7;17;8;6;7;8;16;2;19;17;16;12
    17;1;7;8;3;18;4;17;10;18;15;18;15;2
    10;14;4;8;8;19;13;8;7;13;4;8;13;16
    15;4;8;13;2;7;;1;16;7;18;17;9;2
    12;;2;15;19;8;;12;18;;7;17;;5
    16;;16;7;16;14;;16;19;;13;17;;16
    18;;8;9;2;7;;7;6;;17;;;18
    19;;1;4;5;7;;5;15;;8;;;
    18;;8;10;6;19;;19;8;;9;;;
    14;;16;15;;;;14;13;;;;;
    9;;6;;;;;9;15;;;;;
    

    链接:
    - itertools.zip_longest
    - csv-writer

    【讨论】:

      【解决方案3】:

      假设l 是您的列表,您可以这样做:

      import pandas as pd
      import numpy as np
      pd.DataFrame(np.array(l).reshape(-1,5)).set_index(0).T.to_csv('my_file.csv',index=0)
      

      【讨论】:

      • 这些问题没有明确的答案,因为仍然不清楚。您提供的任何答案都将基于您对 Q 想要什么的假设,而不是 Q 中所述的他的意图
      • 我是这样理解这个问题的。
      • 活了,谢谢 :) 我得出了和你一样的结论(格式明智),但直到他说明他希望发布的答案“思考”是正确的——但在大多数情况下, Q 以这种方式改变三次。
      • 现在可以更好地指定问题并且您的解决方案会中断:) ■
      【解决方案4】:

      如果您的数据集不是太大,您应该先准备好数据,然后将其序列化为 csv 文件。

      import csv
      
      dataset = ['identifier', 1, 2, 3, 4, 'identifier', 10, 11, 12, 13, 'identifier', 21, 22, 23, 24]
      columns = []
      col = []
      for datapoint in dataset:
          if datapoint == 'identifier':
              if col:
                  columns.append(col)
                  col = []
          else:
              col.append(datapoint)
      columns.append(col)
      
      rows_count = max((len(c) for c in columns))
      
      with open('result.csv', 'w') as csvfile:
          writer = csv.writer(csvfile, delimiter=";")
      
          for x in range(rows_count):
              data = []
              for col in columns:
                  if len(col) > x:
                      data.append(col[x])
                  else:
                      data.append("")
              writer.writerow(data)
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2023-01-31
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多