【问题标题】:Compare two CSV files and show only the difference比较两个 CSV 文件并仅显示差异
【发布时间】:2015-03-16 05:23:22
【问题描述】:

我有两个CSV 文件:

文件1.csv

Time, Object_Name, Carrier_Name, Frequency, Longname

2013-08-05 00:00, Alpha, Aircel, 917.86, Aircel_Bhopal

2013-08-05 00:00, Alpha, Aircel, 915.13, Aircel_Indore

文件2.csv

Time, Object_Name, Carrier_Name, Frequency, Longname

2013-08-05 00:00, Alpha, Aircel, 917.86, Aircel_Bhopal

2013-08-05 00:00, Alpha, Aircel, 815.13, Aircel_Indore

这些是示例输入文件,实际上会有很多标题和值,所以我不能对它们进行硬编码。

在我的预期输出中,我想保持前两列和最后一列不变,因为它们不会有任何变化,然后应该对其余的列和值进行比较。

预期输出:

Time, Object_Name, Frequency, Longname

2013-08-05 00:00, 815.13, Aircel_Indore

我该怎么做?

【问题讨论】:

  • 你试过 Linux 中的diff 实用程序吗?
  • 您的问题似乎缺少一些细节。两个文件中的行数总是相同吗?行的顺序可以改变吗(你是否在乎)?固定列(第一、第二、最后)或它们的某些子集是否充当列标识符?最重要的是,如果您有一行在一列中发生变化,而另一行在另一列中发生变化,代码应该输出什么?它应该输出两行的两列,还是应该将未更改的值留空?如果是后者,如何将其与实际变为空的值区分开来?

标签: perl shell


【解决方案1】:

【讨论】:

  • 我已经尝试了所有这些选项,但没有满足我的要求.. :(
  • 我对Perl不熟悉,所以不能给你准确的示例代码。但是也许您可以使用第一个链接中的脚本,第一个答案,而不是推送整行 a - 解析行并仅推送相关字段。解析一行的例子很多。
【解决方案2】:

如果你没有绑定Perl,这里使用AWK的解决方案:

 #!/bin/bash

 awk -v FS="," '

 function filter_columns()
 {
     return sprintf("%s, %s, %s, %s", $1, $2, $(NF-1), $NF);
 }

 NF !=0 && NR == FNR {
    if (NR == 1) {
            print filter_columns();
    } else {
            memory[line++] = filter_columns();
    }
 } NF != 0 && NR != FNR {
    if (FNR == 1) {
            line = 0;
    } else {
            new_line = filter_columns();
            if (new_line != memory[line++]) {
                    print new_line;
            }
    }
 }' File1.csv File2.csv

这个输出:

Time,  Object_Name,  Frequany, Longname
2013-08-05 00:00,  Alpha,  815.13,  Aircel_Indore

这里是解释:

#!/bin/bash

# FS = "," makes awk split each line in fields using
# the comma as separator
awk -v FS="," '

# this function selects the columns you want. NF is the
# the number of field. Therefore $NF is the content of
# the last column and $(NF-1) of the but last.
function filter_columns()
{
     return sprintf("%s, %s, %s, %s", $1, $2, $(NF-1), $NF);
}

# This block processes just the first file, this is the aim
# of the condition NR == FNR. The condition NF != 0 skips the
# empty lines you have in your file. The block prints the header
# and then save all the other lines in the array memory.
NF !=0 && NR == FNR {
    if (NR == 1) {
            print filter_columns();
    } else {
            memory[line++] = filter_columns();
    }
}
# This block processes just the second file (NR != FNR).
# Since the header has been already printed, it skips the first
# line of the second file (FNR == 1). The block compares each line
# against that one saved in the array memory (the corresponding
# line in the first file). The block prints just the lines
# that do not match.
NF != 0 && NR != FNR {
    if (FNR == 1) {
            line = 0;
    } else {
            new_line = filter_columns();
            if (new_line != memory[line++]) {
                    print new_line;
            }
    }
}' File1.csv File2.csv

【讨论】:

    【解决方案3】:

    回答@IlmariKaronen 的问题会更好地澄清问题,但同时我做了一些假设并解决了问题 - 主要是因为我需要一个借口来学习一点 Text::CSV。

    代码如下:

    #!/usr/bin/perl
    
    use strict;
    use warnings;
    
    use Text::CSV;
    use Array::Compare;
    use feature 'say';
    
    open my $in_file, '<', 'infile.csv';
    open my $exp_file, '<', 'expectedfile.csv';
    
    open my $out_diff_file, '>', 'differences.csv';
    
    my $text_csv = Text::CSV->new({ allow_whitespace => 1, auto_diag => 1 });
    
    my $line = readline($in_file);
    my $exp_line = readline($exp_file);
    die 'Different column headers' unless $line eq $exp_line;
    $text_csv->parse($line);
    my @headers = $text_csv->fields();
    
    my %all_differing_indices;
    
    #array-of-array containings lists of "expected" rows for differing lines
    # only columns that differ from the input have values, others are empty
    my @all_differing_rows; 
    
    my $array_comparer = Array::Compare->new(DefFull => 1);
    while (defined($line = readline($in_file))) {
        $exp_line = readline($exp_file);
        if ($line ne $exp_line) {
            $text_csv->parse($line);
            my @in_fields = $text_csv->fields();
            $text_csv->parse($exp_line);
            my @exp_fields = $text_csv->fields();
    
            my @differing_indices = $array_comparer->compare([@in_fields], [@exp_fields]);
            @all_differing_indices{@differing_indices} = (1) x scalar(@differing_indices);
            my @output_row = ('') x scalar(@exp_fields);
            @output_row[0, 1, @differing_indices, $#exp_fields] = @exp_fields[0, 1, @differing_indices, $#exp_fields];
            $all_differing_rows[$#all_differing_rows + 1] = [@output_row];
        }
    }
    
    my @columns_needed = (0, 1, keys(%all_differing_indices), $#headers);
    
    $text_csv->combine(@headers[@columns_needed]);
    say $out_diff_file $text_csv->string();
    for my $row_aref (@all_differing_rows) {
        $text_csv->combine(@{$row_aref}[@columns_needed]);   
        say $out_diff_file $text_csv->string();
    }
    

    它适用于问题中给出的 File1 和 File2 并产生预期的输出(除了 Object_Name 'Alpha' 存在于数据行中 - 我假设这是问题中的错字)。

    Time,Object_Name,Frequany,Longname
    "2013-08-05 00:00",Alpha,815.13,Aircel_Indore
    

    【讨论】:

      【解决方案4】:

      我用非常强大的 linux 工具为它创建了一个脚本。 Link here...

      Linux / Unix - 比较两个 CSV 文件 这个项目是关于两个 csv 文件的比较。

      假设 csvFile1.csv 有 XX 列,而 csvFile2.csv 有 YY 列。

      我编写的脚本应该将 csvFile1.csv 中的一个(键)列与 csvFile2.csv 中的另一个(键)列进行比较。 csvFile1.csv 中的每个变量(键列中的行)将与 csvFile2.csv 中的每个变量进行比较。

      如果 csvFile1.csv 有 1,500 行,而 csvFile2.csv 有 15,000 个,则组合(比较)的总数将为 22,500,000。因此,这是创建可用性报告脚本的非常有用的方法,例如可以将内部产品数据库与外部(供应商)产品数据库进行比较。

      使用的包: csvcut(剪切列) csvdiff(比较两个 csv 文件) ssconvert(将 xlsx 转换为 csv) 图标v curlftpfs 压缩 解压 ntpd proFTPD

      您可以在我的官方博客上找到更多信息(+示例脚本): http://damian1baran.blogspot.sk/2014/01/linux-unix-compare-two-csv-files.html

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2014-06-08
        • 2015-08-10
        • 2017-06-06
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-12-11
        • 2013-06-17
        相关资源
        最近更新 更多