【问题标题】:How can I read lines from the end of file in Perl?如何从 Perl 中的文件末尾读取行?
【发布时间】:2008-11-19 19:28:49
【问题描述】:

我正在编写一个 Perl 脚本来读取 CSV 文件并进行一些计算。 CSV 文件只有两列,如下所示。

One Two
1.00 44.000
3.00 55.000

现在这个 CSV 文件非常大,可以从 10 MB 到 2GB。

目前我正在使用大小为 700 MB 的 CSV 文件。我试图在记事本、excel中打开这个文件,但似乎没有软件可以打开它。

我想读取 CSV 文件中的最后 1000 行并查看值。 我怎样才能做到这一点?我无法在记事本或任何其他程序中打开文件。

如果我编写一个 Perl 脚本,那么我需要处理完整的文件以转到文件末尾,然后读取最后 1000 行。

有没有更好的方法呢?我是 Perl 的新手,任何建议都将不胜感激。

我在网上搜索了一些可用的脚本,例如File::Tail,但我不知道它们是否可以在 Windows 上运行?

【问题讨论】:

    标签: perl large-files


    【解决方案1】:

    File::ReadBackwards 模块允许您以相反的顺序读取文件。只要您不依赖订单,就可以轻松获取最后 N 行。如果您是并且所需的数据足够小(在您的情况下应该如此),您可以将最后 1000 行读入一个数组,然后 reverse 它。

    【讨论】:

    • 第二个推荐。您可以制作自己的搜索/阅读内容,但如果已经在一个广泛使用、经过良好测试的 CPAN 模块中为您完成了这些工作,则毫无意义。
    • 我为想要示例的人编写了一个实现,请参阅gitlab.com/snippets/1957849
    【解决方案2】:

    在*nix中,可以使用tail命令。

    tail -1000 yourfile | perl ...
    

    这只会将最后 1000 行写入 perl 程序。

    在 Windows 上,gnuwin32 和 unxutils 软件包都具有 tail 实用程序。

    【讨论】:

      【解决方案3】:

      这仅与您的主要问题无关,但是当您想检查 File::Tail 等模块是否在您的平台上工作时,请检查来自 CPAN Testers 的结果。 CPAN Search 模块页面顶部的链接将引导您到


      (来源:flickr.com)

      查看矩阵,您会发现在 Windows 上测试的所有 Perl 版本上,该模块确实存在问题:


      (来源:flickr.com)

      【讨论】:

      • 现在检查并且 File::tail 1.3 正在 Windows 上传递。
      【解决方案4】:

      没有tail,仅Perl的解决方案并不是那么不合理。

      一种方法是从文件末尾查找,然后从中读取行。如果你没有足够的行数,请从末尾再往前寻找,然后再试一次。

      sub last_x_lines {
          my ($filename, $lineswanted) = @_;
          my ($line, $filesize, $seekpos, $numread, @lines);
      
          open F, $filename or die "Can't read $filename: $!\n";
      
          $filesize = -s $filename;
          $seekpos = 50 * $lineswanted;
          $numread = 0;
      
          while ($numread < $lineswanted) {
              @lines = ();
              $numread = 0;
              seek(F, $filesize - $seekpos, 0);
              <F> if $seekpos < $filesize; # Discard probably fragmentary line
              while (defined($line = <F>)) {
                  push @lines, $line;
                  shift @lines if ++$numread > $lineswanted;
              }
              if ($numread < $lineswanted) {
                  # We didn't get enough lines. Double the amount of space to read from next time.
                  if ($seekpos >= $filesize) {
                      die "There aren't even $lineswanted lines in $filename - I got $numread\n";
                  }
                  $seekpos *= 2;
                  $seekpos = $filesize if $seekpos >= $filesize;
              }
          }
          close F;
          return @lines;
      }
      

      附:更好的标题应该是“从 Perl 中的大文件末尾读取行”。

      【讨论】:

      • P.P.S.添加评论以解释否决票将不胜感激。如果我觉得答案没有帮助/没有回应,我会删除它。
      • 同意,这是一个相当有效的解决方案和有用的信息。 +1。
      【解决方案5】:

      我在纯 Perl 上使用以下代码编写了快速向后文件搜索:

      #!/usr/bin/perl 
      use warnings;
      use strict;
      my ($file, $num_of_lines) = @ARGV;
      
      my $count = 0;
      my $filesize = -s $file; # filesize used to control reaching the start of file while reading it backward
      my $offset = -2; # skip two last characters: \n and ^Z in the end of file
      
      open F, $file or die "Can't read $file: $!\n";
      
      while (abs($offset) < $filesize) {
          my $line = "";
          # we need to check the start of the file for seek in mode "2" 
          # as it continues to output data in revers order even when out of file range reached
          while (abs($offset) < $filesize) {
              seek F, $offset, 2;     # because of negative $offset & "2" - it will seek backward
              $offset -= 1;           # move back the counter
              my $char = getc F;
              last if $char eq "\n"; # catch the whole line if reached
              $line = $char . $line; # otherwise we have next character for current line
          }
      
          # got the next line!
          print $line, "\n";
      
          # exit the loop if we are done
          $count++;
          last if $count > $num_of_lines;
      }
      

      然后像这样运行这个脚本:

      $ get-x-lines-from-end.pl ./myhugefile.log 200
      

      【讨论】:

      • 我发布了这个因为我不喜欢这个 "$seekpos *= 2;"以前的解决方案中的方法也是通过充分工作来实现的
      【解决方案6】:
      perl -n -e "shift @d if (@d >= 1000); push(@d, $_); END { print @d }" < bigfile.csv
      

      虽然说真的,UNIX 系统可以简单地使用tail -n 1000 的事实应该说服您简单地安装cygwin 或colinux

      【讨论】:

        【解决方案7】:

        我相信你可以使用 Tie::File 模块。看起来这会将行加载到数组中,然后您可以获取数组的大小并将 arrayS-ze-1000 处理为 arraySize-1。

        Tie::File

        另一个选项是计算文件中的行数,然后遍历文件一次,并开始读取 numberofLines-1000 处的值

        $count = `wc -l < $file`;
        die "wc failed: $?" if $?;
        chomp($count);
        

        这会给你行数(在大多数系统上。

        【讨论】:

          【解决方案8】:

          如果你知道文件的行数,你可以这样做

          perl -ne "print if ($. > N);" filename.csv
          

          其中 N 是 $num_lines_in_file - $num_lines_to_print。 你可以用

          来计算行数
          perl -e "while (<>) {} print $.;" filename.csv
          

          【讨论】:

            【解决方案9】:

            模块是要走的路。但是,有时您可能正在编写一段代码,希望在各种机器上运行,而这些机器可能缺少更晦涩的 CPAN 模块。在这种情况下,为什么不只是'tail'并将输出从 Perl 中转储到临时文件中?

            #!/usr/bin/perl
            
            `tail --lines=1000 /path/myfile.txt > tempfile.txt`
            

            如果安装一个 CPAN 模块可能会出现问题,那么您就有了一些不依赖于 CPAN 模块的东西。

            【讨论】:

            • 他确实说过记事本和 Excel,/usr/bin/perl + tail 可能不太奏效。
            【解决方案10】:

            如果不依赖 tail,我可能会这样做,如果你有超过 $FILESIZE [2GB?] 的内存,那么我会偷懒做:

            my @lines = <>;
            my @lastKlines = @lines[-1000,-1];
            

            尽管涉及tail 或seek() 的其他答案几乎可以解决这个问题。

            【讨论】:

            • 好吧,使用tail,就像你不知道一样。您询问了 perl,这在 perl 中有效。如果有任何理由应将其视为不恰当的答案,我将不胜感激。
            【解决方案11】:

            您绝对应该使用 File::Tail,或者更好的另一个模块。它不是一个脚本,它是一个模块(编程库)。它可能适用于 Windows。正如有人所说,您可以在 CPAN Testers 上进行检查,或者通常只需阅读模块文档或尝试一下即可。

            您选择使用 tail 实用程序作为首选答案,但在 Windows 上这可能比 File::Tail 更令人头疼。

            【讨论】:

              猜你喜欢
              • 2023-03-12
              • 1970-01-01
              • 2011-09-21
              • 1970-01-01
              • 2011-08-07
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 2013-03-27
              相关资源
              最近更新 更多