【问题标题】:perl extract text between SAME delimiter using flip-flopperl 使用触发器在 SAME 分隔符之间提取文本
【发布时间】:2017-02-01 00:59:25
【问题描述】:

过去,我已经能够使用触发器来提取具有不同 START 和 END 的文本。 这次我在尝试提取文本时遇到了很多麻烦,因为我的源文件中没有不同的分隔符,因为触发器的 START 和 END 是相同的。我希望触发器在年份为 yyyy 的行中启动并继续将$_ 推送到数组,直到另一行开始 yyyy。 触发器的问题是它会在我的下一个 START 时为假。

while (<SOURCEFILE>) {
  print if (/^2017/ ... /^2017/) 
}

对给定的源数据使用上述内容将错过我还需要匹配的文件的第二个多行部分。也许我认为是解析多行文件的最佳方法的触发器在这种情况下不起作用?我想要做的是从以日期开头的第一行开始匹配并继续匹配,直到以日期开头的下一行之前的行。

样本数据是:

2017 message 1
Text
Text

Text

2017 message 2
more text
more text

more text

2017 message 3
yet more text
yet more text

yet more text

但我得到了:

2017 message 1
Text
Text

Text

2017 message 2
2017 message 3
yet more text
yet more text

yet more text

...缺少消息 2 内容..

我不能在源数据中依赖空格或不同的 END 分隔符。 我想要的是打印每条消息(实际上是push @myarray, $_ & 然后测试匹配),但是这里我缺少消息 2 下面的行,因为触发器设置为 false。有什么方法可以用触发器来处理这个问题,还是我需要使用其他东西? 提前感谢任何可以提供帮助/建议的人。

【问题讨论】:

  • 我遇到了一个类似的问题,其中包含由相同 SQL 注释分隔的块的 SQL 文件。当我未能使用触发器运算符时,我最终将整个文件读入带有perl -ne "BEGIN { $/ = undef } ..." 的变量中,并使用多行正则表达式来匹配所需的块。

标签: perl delimiter flip-flop


【解决方案1】:

这里有一个方法:

use Modern::Perl;
use Data::Dumper;
my $part = -1;
my $parts;
while(<DATA>) {
    chomp;
    if (/^2017/ .. 1==0) {
        $part++ if /^2017/;
        push @{$parts->[$part]}, $_;
    }
}
say Dumper$parts;

__DATA__
2017 message 1
Text
Text

Text

2017 message 2
more text
more text

more text

2017 message 3
yet more text
yet more text

yet more text

输出:

$VAR1 = [
          [
            '2017 message 1',
            'Text',
            'Text',
            '',
            'Text',
            ''
          ],
          [
            '2017 message 2',
            'more text',
            'more text',
            '',
            'more text',
            ''
          ],
          [
            '2017 message 3',
            'yet more text',
            'yet more text',
            '',
            'yet more text'
          ]
        ];

【讨论】:

  • 嗯,很好看。 if 行实际上是多余的,不是吗?
  • @Sobrique,用于跳过第一个匹配/^2017/之前的行。
【解决方案2】:

正如您所发现的,带有匹配分隔符的触发器不能很好地工作。

您是否考虑改为设置$/?

例如:

#!/usr/bin/env perl
use strict;
use warnings; 

local $/ = "2017 message";
my $count;

while ( <DATA> ) {

    print "\nStart of block:", ++$count, "\n";

    print;

    print "\nEnd of block:", $count, "\n";
}

__DATA__
2017 message 1
Text
Text

Text

2017 message 2
more text
more text

more text

2017 message 3
yet more text
yet more text

yet more text

虽然它并不完美,因为它在分隔符上分割文件 - 这意味着在第一个之前有一个“位”(所以你得到 4 个块)。您可以通过明智地使用 'chomp' 来重新拼接它,这会从当前块中删除 $/:

#!/usr/bin/env perl
use strict;
use warnings; 

local $/ = "2017 message";
my $count;

while ( <DATA> ) {
    #remove '2017 message'
    chomp;
    #check for empty (first) block
    next unless /\S/;
    print "\nStart of block:", ++$count, "\n";
    #re add '2017 message'
    print $/;
    print;

    print "\nEnd of block:", $count, "\n";
}

或者,数组数组怎么样,每次点击消息时更新“目标键”?

#!/usr/bin/env perl
use strict;
use warnings; 

use Data::Dumper;

my %messages; 
my $message_id;
while ( <DATA> ) {
   chomp;
   if ( m/2017 message (\d+)/ ) { $message_id = $1 }; 
   push @{ $messages{$message_id} }, $_; 
}

print Dumper \%messages;

注意 - 我使用的是散列,而不是数组,因为对于不从零连续开始的消息排序,这更可靠。 (并且使用这种方法的数组将有一个空的“第零”元素)。

注意 - 它也将有“空”'' 元素供您使用空白行。如果你愿意,你可以过滤这些。

【讨论】:

    【解决方案3】:

    我不知道如何用触发器做到这一点。一年前我试过了。但是我用一些逻辑做了同样的事情。

    my $line_concat;
    my $f = 0;
    while (<DATA>) {
        if(/^2017/ && !$f) {
            $f = 1;
        }
    
        if (/^2017/) {
            print "$line_concat\n" if $line_concat ne "";
            $line_concat = "";
        }
    
        $line_concat .= $_ if $f;
    }
    
    print $line_concat if $line_concat ne "";
    

    【讨论】:

    • 还有一件事......它有效,但你永远不会重置$f。为什么不?您可以将 reset 和 redo 放入第二个 if 块中,以使每一对实际上都是一对,但我认为这样做没有好处。
    【解决方案4】:

    您只需要一个缓冲区来累积行,直到找到匹配的 /^20\d\d[ ]/ 或文件结尾。

    my $in = 0;
    my @buf;
    while (<>) {
       if ($in && /^20\d\d[ ]/) {
          process(@buf);
          @buf = ();
          $in = 0;
       }
    
       push @buf, $_ if $in ||= /^2017[ ]/;
    }
    
    process(@buf) if $in;
    

    我们可以重新排列代码以使其仅在一个位置处理记录,从而允许内联process。

    my $in = 0;
    my @buf;
    while (1) {
       $_ = <>;
    
       if ($in && (!defined($_) || /^20\d\d[ ]/)) {
          process(@buf);
          @buf = ();
          $in = 0;
       }
    
       last if !defined($_);
    
       push @buf, $_ if $in ||= /^2017[ ]/;
    }
    

    【讨论】:

    • 上述 cmets 从来都不是正确的,并且已因答案的更改而过时。
    猜你喜欢
    • 2016-01-03
    • 1970-01-01
    • 1970-01-01
    • 2016-02-10
    • 2021-04-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多