【问题标题】:Merging CSV lines in huge file合并大文件中的 CSV 行
【发布时间】:2015-07-09 13:26:37
【问题描述】:

我有一个像这样的 CSV

783582893T,2014-01-01 00:00,0,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582893T,2014-01-01 00:15,1,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582893T,2014-01-01 00:30,2,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582855T,2014-01-01 00:00,0,128,35.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582855T,2014-01-01 00:15,1,128,35.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582855T,2014-01-01 00:30,2,128,35.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
...
783582893T,2014-01-02 00:00,0,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582893T,2014-01-02 00:15,1,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582893T,2014-01-02 00:30,2,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y

虽然有 50 亿条记录。如果您注意到第一列和第二列(当天)的一部分,则其中三个记录都“分组”在一起,只是当天前 30 分钟的 15 分钟间隔的细分。

我希望输出看起来像

783582893T,2014-01-01 00:00,0,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
783582855T,2014-01-01 00:00,0,128,35.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y
...
783582893T,2014-01-02 00:00,0,124,29.1,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y,40.0,0.0,40,40,5,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,40,Y

重复行的前 4 列被省略,其余列与同类的第一条记录相结合。基本上我将一天从每行 15 分钟转换为每行 1 天。

由于我将处理 50 亿条记录,我认为最好的办法是使用正则表达式(和 EmEditor)或为此而设计的一些工具(多线程、优化),而不是自定义编程解决方案。虽然我对 nodeJS 或 C# 中相对简单且超级快速的想法持开放态度。

如何做到这一点?

【问题讨论】:

  • 每条记录总是有 3 条记录吗?
  • 因此,按第一列对行进行分组,从第一行获取前 4 列(组中的所有行都应该相同),然后从其他行到第一行。对吗?
  • 是的,总是有 3 条记录(技术上是 96 条记录,但适用于这 3 条记录的解决方案应该不难适应 96 条记录)
  • 而需要组合在一起的行(具有相同的第一列)总是连续的?它们永远不可能出现故障?
  • 是的,它们不会出现故障。我编辑了一下没有任何价值,因为从技术上讲,您是在第一列和第二列的一部分进行分组。正如我所说,它从每行 15 分钟的间隔变为每行 1 天。因此,如果您愿意,日期时间的日期部分也是分组“键”的一部分。

标签: c# regex node.js csv sed


【解决方案1】:

如果总是有一定数量的记录记录并且它们是有序的,那么一次读取几行并解析和输出它们会相当容易。尝试对数十亿条记录进行正则表达式将花费很长时间。使用StreamReader 和StreamWriter 应该可以读取和写入这些大文件,因为它们一次读取和写入一行。

using (StreamReader sr = new StreamReader("inputFile.txt")) 
using (StreamWriter sw = new StreamWriter("outputFile.txt"))
{
    string line1;
    int counter = 0;
    var lineCountToGroup = 3; //change to 96
    while ((line1 = sr.ReadLine()) != null) 
    {
        var lines = new List<string>();
        lines.Add(line1);
        for(int i = 0; i < lineCountToGroup - 1; i++) //less 1 because we already added line1
            lines.Add(sr.ReadLine());

        var groupedLine = lines.SomeLinqIfNecessary();//whatever your grouping logic is
        sw.WriteLine(groupedLine);
    }
}

免责声明 - 未经测试的代码,没有错误处理,并假设重复的行数确实正确,等等。您显然需要针对您的确切场景进行一些调整。

【讨论】:

  • 我怀疑这对于 50 亿条记录来说是一个很好的解决方案。为什么不按顺序读/写?
  • 文件大小为 500 GB,我认为这会将其全部读入内存。
  • 应该不难适应通过逐行加载和保存而不是一次全部工作的东西。
  • @ParoX 非常正确。我已经编辑了我的答案以使用 StreamReader 和 StreamWriter 一次读取和写入一行,我认为这会起作用
  • 是的。但我认为 OP 必须写 ReadLine 96 次。根据此评论:是的,总是有 3 条记录(从技术上讲,有 96 条记录,但适用于这 3 条记录的解决方案应该不难适应 96 条记录)最好将其放入循环中
【解决方案2】:

你可以做这样的事情(未经测试的代码没有任何错误处理 - 但应该给你它的一般要点):

using (var sin = new SteamReader("yourfile.csv")
using (var sout = new SteamWriter("outfile.csv")
{
    var line = sin.ReadLine();    // note: should add error handling for empty files
    var cells = line.Split(",");  // note: you should probably check the length too!
    var key = cells[0];           // use this to match other rows
    StringBuilder output = new StringBuilder(line);   // this is the output line we build
    while ((line = sin.ReadLine()) != null) // if we have more lines
    {
        cells = line.Split(",");    // split so we can get the first column
        while(cells[0] == key)      // if the first column matches the current key
        {
            output.Append(String.Join(",",cells.Skip(4)));   // add this row to our output line
        }
        // once the key changes
        sout.WriteLine(output.ToString());      // write out the line we've built up
        output.Clear();
        output.Append(line);         // update the new line to build
        key = cells[0];              // and update the key
    }
    // once all lines have been processed
    sout.WriteLine(output.ToString());    // We'll have just the last line to write out
}

这个想法是依次循环遍历每一行并跟踪第一列的当前值。当该值更改时,您写出您一直在构建的output 行并更新key。这样您就不必担心您有多少匹配项,或者您是否可能会丢失一些分数。

请注意,如果要连接 96 行,使用 StringBuilder 表示 output 可能比 String 更有效。

【讨论】:

    【解决方案3】:

    定义 ProcessOutputLine 以存储合并的行。 在每个 ReadLine 之后和文件末尾调用 ProcessLine。

    string curKey     =""   ; 
    string keyLength  = ... ; // set totalength of 4 first columns
    string outputLine = ""  ;
    
    private void ProcessInputLine(string line)
    {
      string newKey=line.substring(0,keyLength) ;
      if (newKey==curKey) outputline+=line.substring(keyLength) ;
      else 
      { 
        if (outputline!="") ProcessOutPutLine(outputLine)
        curkey = newKey ;
        outputLine=Line ;
    }
    

    编辑:这个解决方案与 Matt Burland 的解决方案非常相似,唯一明显的区别是我不使用拆分功能。

    【讨论】:

    • 实际上一个非常重要的区别是您依赖于列的固定宽度,如果它们实际上是固定的,这很好,但如果不是,则没有那么多。现在可能(尽管 OP 需要确认)可以安全地猜测 至少第一列是固定长度的。
    猜你喜欢
    • 2021-05-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-04-19
    • 1970-01-01
    • 1970-01-01
    • 2015-04-01
    • 1970-01-01
    相关资源
    最近更新 更多