【问题标题】:Remove/Extract rows based on Unique/duplicate Id from a CSV file根据 CSV 文件中的唯一/重复 ID 删除/提取行
【发布时间】:2015-10-14 17:18:35
【问题描述】:

根据您的看法,我需要根据 Id 是否唯一来删除行或如果 Id 有重复项则提取行(保留所有重复项)。 而且我不确定/没有足够的 Perl 知识来完成这个。我找到了类似的主题,但没有太多成功。这些是我使用example 1、example 2 和example 3 的示例。在上一个问题中,有人向我展示了 List::MoreUtils 模块的解决方案,因此我可以将值与通用 ID 合并。现在不是这种情况,如果 id 是唯一的,这将删除行。我知道我可能可以使用 List::MoreUtils 模块来做到这一点,但我想不这样做。这是我的虚拟数据(从其他问题复制的示例数据,因为数据无关紧要),在这里你可以看到我所追求的。顺序并不重要。

之前:

Cat_id;Cat_name;Id;Name;Amount;Colour;Bla
101;Fruits;50010;Grape;500;Red;1
101;Fruits;50020;Strawberry;500;Red;1
201;Vegetables;60010;Carrot;500;White;1
101;Fruits;50060;Apple;1000;Red;1
101;Fruits;50030;Banana;1000;Green;1
101;Fruits;50060;Apple;500;Green;1
101;Fruits;50020;Strawberry;1000;Red;1
201;Vegetables;60010;Carrot;100;Purple;1
101;Fruits;50020;Strawberry;200;Red;1

之后:

Cat_id;Cat_name;Id;Name;Amount;Colour;Bla
101;Fruits;50020;Strawberry;500;Red;1
201;Vegetables;60010;Carrot;500;White;1
101;Fruits;50060;Apple;1000;Red;1
101;Fruits;50060;Apple;500;Green;1
101;Fruits;50020;Strawberry;1000;Red;1
201;Vegetables;60010;Carrot;100;Purple;1
101;Fruits;50020;Strawberry;200;Red;1

您可以看到 ID 为 50010 和 50030 的 Grape 和 Banana 行已被删除,因为两者只存在一个条目。

这是我的脚本,我正在努力从哈希中选择唯一值并输出它们的部分(考虑到 Text::CSV_XS 模块)。有人可以告诉我怎么做吗?

#!/usr/bin/perl -w
use strict;
use warnings;
use Text::CSV_XS;

my $inputfile = shift || die "Give input and output names!\n";
my $outputfile = shift || die "Give output name!\n";

open (my $infile, '<:encoding(iso-8859-1)', $inputfile) or die "Sourcefile in use / not found :$!\n";
open (my $outfile, '>:encoding(UTF-8)', $outputfile) or die "Outputfile in use :$!\n";

my $csv_in = Text::CSV_XS->new({binary => 1,sep_char => ";",auto_diag => 1,always_quote => 1,eol => $/}); 
my $csv_out = Text::CSV_XS->new({binary => 1,sep_char => "|",auto_diag => 1,always_quote => 1,eol => $/});

my $header = $csv_in->getline($infile);
$csv_out->print($outfile, $header);

my %data;

while (my $elements = $csv_in->getline($infile)){
    my @columns = @{ $elements };       
    my $id = $columns[2];
    push @{ $data{$id} }, \@columns;
}

for my $id ( sort keys %data ){                 # Sort not important
    if @{ $data{$id} } > 1                      # Here I have no idea anymore..
        $csv_out->print($outfile, \@columns);   #
}

【问题讨论】:

  • 这个问题看起来有点眼熟。 stackoverflow.com/questions/28627669/…
  • @Sobrique 同意,几乎一样.. 我试图解决那个问题,但是如果 id 相同,那就是合并字段,如果 id 是唯一的,这个是删除行

标签: perl csv


【解决方案1】:

与其用整个数据集加载哈希,我想我会继续读取文件两次,只用你的ID 值加载一个哈希。这肯定会花费更长的时间,但随着文件的增长,将所有数据都存储在内存中可能会带来不利影响。

也就是说,我没有使用Text::CSV_XS,但这是我的概念想法。

my %count;

open (my $infile, '<:encoding(iso-8859-1)', $inputfile) or die;
open (my $outfile, '>:encoding(UTF-8)', $outputfile) or die;

while (<$infile>) {
  next if $. == 1;
  my ($id) = (split /;/, $_, 4)[2];
  $count{$id}++;
}

seek $infile, 0, 0;

while (<$infile>) {
  my @fields = split /;/;
  print $outfile join '|', @fields if $count{$fields[2]} > 1 or $. == 1;    
}

close $infile;
close $outfile;

末尾的$. == 1 是为了避免丢失标题行。

-- 编辑--

#!/usr/bin/perl -w

use strict;
use warnings;
use Text::CSV_XS;

my $inputfile = shift || die "Give input and output names!\n";
my $outputfile = shift || die "Give output name!\n";

open (my $infile, '<:encoding(iso-8859-1)', $inputfile) or die;
open (my $outfile, '>:encoding(UTF-8)', $outputfile) or die;

my $csv_in = Text::CSV_XS->new({binary => 1,sep_char => ";",
    auto_diag => 1,always_quote => 1,eol => $/}); 
my $csv_out = Text::CSV_XS->new({binary => 1,sep_char => "|",
    auto_diag => 1,always_quote => 1,eol => $/});

my ($count, %count) = (1);

while (my $elements = $csv_in->getline($infile)){
  $count{$$elements[2]}++;
}

seek $infile, 0, 0;

while (my $elements = $csv_in->getline($infile)){
  $csv_out->print($outfile, $elements)
    if $count{$$elements[2]} > 1 or $count++ == 1;
}

close $infile;
close $outfile;

【讨论】:

  • 感谢您的回答,但我必须使用 Text::CSV_XS 模块(它是一个很大的文件,数据中有分隔符)。您对如何使用该模块有什么建议吗?
  • 我认为您的其余代码都很好......我只是懒得在概念上描述我将如何做到这一点。您可以完全按照您的方式采用现有的 Text::CSV_XS。我已经为此修改了我的回复。
  • 感谢您的编辑,但现在它显示“在第 27 行的哈希元素中使用未初始化的值”,如:它没有可打印的内容。我错过了什么?
  • 你好 Jan。我猜没有看到你的数据,但我最初的猜测是你的第一个文件 ($infile) 中有一行是空白的,或者没有>= 四列数据,这使得$$elements[2]。您可以打印到标准输出(以调试)以查看它在出现此错误之前得到了多少行?或者,只是为了调试,公开该行上的每个变量以查看它们在引发错误时包含的内容(即print $elements $csv_out $outfile\n)。我认为可以肯定地说 %count 有数据。
猜你喜欢
  • 2017-10-22
  • 2014-06-30
  • 2012-02-20
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-10-18
  • 2018-09-29
  • 2020-03-26
相关资源
最近更新 更多