【问题标题】:Perl - Changing file name in the middle of writePerl - 在写入过程中更改文件名
【发布时间】:2015-11-25 18:12:11
【问题描述】:

我正在尝试获取我在 Perl 中创建的一个非常大的 txt 文件(超过一百万行)并通过 Perl 中的不同语句运行它,该语句基本上看起来像这样(注意以下是 shell)

a=0
b=1
while read line;
do
    echo -n "" > "Write file"${b}
    a=($a + 1)
    while ( $a <= 5000)
    do
        echo $line >> "Write file"${b}
        a=($a + 1)
    done
    a=0
    b=($b + 1)
done < "read file"

尝试将其大小缩小到每个文件 5k 行,并且每次递增(filename1.txt、filename2.txt、filename3.txt 等)
这似乎在 shell 中不起作用,可能是由于输入文件的大小,对于我的生活,我想不出如何在循环中间更改我正在写入的文件..

【问题讨论】:

    标签: perl split increment large-files filesplitting


    【解决方案1】:

    您可以在 shell 中使用 split 执行此操作。

    例如:

    split -l 5000 filename.txt filename.txt.
    

    将filename.txt 拆分为多个文件,每个文件最多5,000 行。输出文件将是名称filename.txt.aa、filename.txt.ab、filename.txt.ac 等。

    来自我的man split:

    NAME
         split -- split a file into pieces
    
    SYNOPSIS
         split [-a suffix_length] [-b byte_count[k|m]] [-l line_count] [-p pattern] [file [name]]
    
    DESCRIPTION
         The split utility reads the given file and breaks it up into files of 1000 lines each.  If file is a single dash (`-') or absent, split reads from the stan-
         dard input.
    
         The options are as follows:
    
         -a suffix_length
                 Use suffix_length letters to form the suffix of the file name.
    
         -b byte_count[k|m]
                 Create smaller files byte_count bytes in length.  If ``k'' is appended to the number, the file is split into byte_count kilobyte pieces.  If ``m'' is
                 appended to the number, the file is split into byte_count megabyte pieces.
    
         -l line_count
                 Create smaller files n lines in length.
    
         -p pattern
                 The file is split whenever an input line matches pattern, which is interpreted as an extended regular expression.  The matching line will be the
                 first line of the next output file.  This option is incompatible with the -b and -l options.
    
         If additional arguments are specified, the first is used as the name of the input file which is to be split.  If a second additional argument is specified,
         it is used as a prefix for the names of the files into which the file is split.  In this case, each file into which the file is split is named by the prefix
         followed by a lexically ordered suffix using suffix_length characters in the range ``a-z''.  If -a is not specified, two letters are used as the suffix.
    
         If the name argument is not specified, the file is split into lexically ordered files named with the prefix ``x'' and with suffixes as above.
    

    【讨论】:

      【解决方案2】:

      顺便说一句,这是你的固定脚本:

      #!/bin/sh
      a=0
      b=1
      while read line; do
          if [ $a -eq 0 ]; then
              echo -n '' > out-file-${b}
          fi
      
          echo $line >> out-file-${b}
      
          a=$(( $a + 1 ))
          if [ $a -eq 10 ]; then
              a=0
              b=$(( $b + 1 ))
          fi
      done < in-file
      

      用bash 和dash 测试。

      【讨论】:

      • 我很惊讶这被接受为答案。使用split 会好得多。如果可以的话,我会将此作为评论发布,因为它旨在帮助您在未来的努力中,而不是这种特殊情况。
      • 因为当我们谈论由单个“拆分”产生的数百个文件时,拆分对于文件命名来说太混乱了。我需要一个增量计数器来保持整洁。
      • 这与split 的做法有何不同?
      • 拆分后会变成 file.txt.aa、file.txt.ab 等等。用于文件传输和自动读取到不同团队的 unix 框和无法飞行的表中。至少对于我正在为之制作的团队而言……他们想要 file0001.txt、file0002.txt 等等,所以这就是为什么拆分对我不起作用的原因,否则我会首先使用它
      猜你喜欢
      • 1970-01-01
      • 2012-09-26
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-02-06
      • 1970-01-01
      • 1970-01-01
      • 2011-12-31
      相关资源
      最近更新 更多