【问题标题】:How to find text in data file and calculate average using perl如何在数据文件中查找文本并使用 perl 计算平均值
【发布时间】:2014-03-10 12:22:13
【问题描述】:

我想替换一个 grep |哇 | perl 命令和纯 perl 解决方案,使其运行起来更快更简单。

我想将 input.txt 中的每一行与一个 data.txt 文件进行匹配,并计算匹配 ID 名称和数字的值的平均值。

input.txt 包含 1 列 ID 号:

FBgn0260798
FBgn0040007
FBgn0046692

我想将每个 ID 号与其对应的 ID 名称和相关值进行匹配。这是 data.txt 的示例,其中第 1 列是 ID 号,第 2 列和第 3 列是 ID name1 和 ID name2,第 3 列包含我要计算平均值的值。

FBgn0260798 CG17665 CG17665 21.4497
FBgn0040007 Gprk1   CG40129 22.4236
FBgn0046692 RpL38   CG18001 1182.88

到目前为止,我使用 grep 和 awk 生成了一个包含匹配 ID 编号和值的相应值的输出文件,然后使用该输出文件使用以下命令计算计数和平均值:

# First part using grep | awk
exec < input.txt
while read line
    do
            grep -w $line data.txt | cut -f1,2,3,4 | awk '{print $1,$2,$3,$4} ' >> output.txt
    done
 # Second part with perl

open my $input, '<', "output_1.txt" or die; ## the output file is from the first part and has the same layout as the data.txt file

my $total = 0;
my $count = 0;

while (<$input>) {

    my ($name, $id1, $id2, $value) = split;
    $total += $value;
    $count += 1;

}

print "The total is $total\n";
print "The count is $count\n";
print "The average is ", $total / $count, "\n";

这两个部分都可以正常工作,但我想通过只运行一个脚本来简化它。我一直在尝试找到一种更快的方法来在 perl 中一起运行所有内容,但是经过几个小时的阅读,我完全不知道该怎么做。我一直在玩散列、数组、if 和 elsif 语句,但成功率为零。如果有人有建议等,那就太好了。

谢谢, 哈丽特

【问题讨论】:

  • 那么,上面的 perl 可以正常工作,但是速度太慢了?
  • 请解释您的要求。您显示的程序将成功打印output_1.txt 第四列的平均值。你还需要吗?您似乎要求使用纯 Perl 解决方案来替换 grep | awk | perl 命令。请更彻底地解释,并显示您的完整命令行。

标签: perl if-statement awk grep


【解决方案1】:

如果我理解你的话,你有一个数据文件,其中包含每一行的 name 和该行的 value。其他两个 ID 并不重要。

您将使用一个名为输入文件的新文件,该文件将包含在数据文件中找到的匹配名称。这些是您想要平均的值。

最快的方法是创建一个以 names 为键的散列,值将是该 name 中的 value 数据文件。因为这是一个哈希,所以可以快速定位到对应的值。这比一遍又一遍地 grep 同一个数组要快得多。

这第一部分将读取data.txt 文件并将name 和value 存储在由name 键入的哈希中。

use strict;
use warnings;
use autodie;   # This way, you don't have to check if you can't open the file
use feature qw(say);

use constant {
    INPUT_NAME  => "input.txt",
    DATA_FILE   => "data.txt",
};

#
# Read in data.txt and get the values and keys
#
open my $data_fh, "<", DATA_FILE;
my %ids;
while ( my $line = <$data_fh> ) {
    chomp $line;
    my ($name, $id1, $id2, $value) = split /\s+/, $line;
    $ids{$name} = $value;
}
close $data_fh;

现在,您有了这个哈希,就很容易阅读input.txt 文件并在data.txt 文件中找到匹配的名称:

open $input_fh, "<", INPUT_FILE;
my $count = 0;
my $total = 0;
while ( my $name = <$input_fh> ) {
    chomp $name;
    if ( not defined $ids{$name} ) {
         die qq(Cannot find matching id "$name" in data file\n);
    }
    $total += $ids{$name};
    $count += 1;
}
close $input_fh;
say "Average = " $total / $count;

您通读每个文件一次。我假设每个文件中每个 name 只有一个实例。

【讨论】:

  • 感谢您的解释,它非常详细,有助于我的理解(我最近开始学习 Perl)。一个问题,我不断收到此错误消息:match_avg_test.pl 第 10 行的语法错误,靠近“DATA_FILE” test.pl 由于编译错误而中止。我已经注释掉了严格的警告以消除显式名称错误消息并检查了所有内容但我看不到它。如果您有任何想法,那就太好了
  • 这些行应该以逗号而不是分号结尾。对不起。
猜你喜欢
  • 2012-12-06
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-06-30
相关资源
最近更新 更多