【问题标题】:How to extract lines from a column separated by empty lines?如何从由空行分隔的列中提取行?
【发布时间】:2017-01-13 22:45:14
【问题描述】:

我有一个大的标记句子文件。不同的句子由空行隔开。输入文件基本上只是一大列。

我想转置单列,使每个唯一的句子都有自己的行。

输入:

Sentence1
Sentence1
Sentence1
Sentence1

Sentence2
Sentence2
Sentence2

...

SentenceN

期望的输出是这样的:

Sentence1 Sentence1 Sentence1 Sentence1
Sentence2 Sentence2 Sentence2
...

我一直在寻找 grep、awk、sed 和 tr,但我正在努力使用正确的语法。

谢谢!

【问题讨论】:

    标签: unix text-processing


    【解决方案1】:

    如果您明智地选择记录和字段分隔符,awk 很简单:

    awk '$1=$1' RS= FS="\n" OFS=" " infile
    

    输出:

    Sentence1 Sentence1 Sentence1 Sentence1
    Sentence2 Sentence2 Sentence2
    ...
    SentenceN
    

    说明

    • RS= 将记录分隔符设置为“空行”。
    • FS="\n" 将字段分隔符设置为换行符。
    • OFS=" " 将输出分隔符设置为空格。
    • $1=$1 重新评估输入并根据FS 对其进行拆分。这也评估为真,因此以OFS 作为分隔符输出输入。

    【讨论】:

      【解决方案2】:

      perl 很容易:

      #!/usr/bin/env perl
      
      use strict;
      use warnings;
      
      local $/ = "\n\n";
      
      while ( <DATA> ) {
         s/\n/ /g;
         print;
         print "\n";
      }
      
      __DATA__
      Sentence1
      Sentence1
      Sentence1
      Sentence1
      
      Sentence2
      Sentence2
      Sentence2
      

      或单线化:

      perl -00 -pe 's/\n/ /g' 
      

      【讨论】:

      • 感谢您的帮助。我当然离我的目标更近了,但是这些句子现在在一行中。我的目标是让一个句子成行。
      • @seppo:它需要-l (dash-ell) 在记录后添加换行符。此外,您可能希望在替换命令后添加 $\ = "\n" 以减少记录之间的间距。
      【解决方案3】:

      awk 解决方案

      awk '{ if($1~"^$") {print a;a="";} else a=a" "$0;} END {print a}' test.txt
      

      【讨论】:

        猜你喜欢
        • 2023-03-25
        • 1970-01-01
        • 2017-09-09
        • 2016-07-08
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多