【问题标题】:parse multiple lines in perl regular expression and extract value解析perl正则表达式中的多行并提取值
【发布时间】:2010-12-15 14:25:15
【问题描述】:

我是 perl 的初学者。我有一个文本文件,其文本类似于以下内容。我需要提取 VALUE="NEEDED VALUE>"。说菠菜,我应该单独吃沙拉。

如何使用 perl 正则表达式来获取值。我需要解析多行才能得到它。即在每个#ifonly --- #endifonly

之间

$cat check.txt

while (<$file>)
{
   if (m/#ifonly .+ SPINACH .+ VALUE=(")([\w]*)(") .+ #endifonly/g)
{
    my $chosen = $2;
   }
}


#ifonly APPLE CARROT SPINACH
VALUE="SALAD" REQUIRED="yes" 
QW RETEWRT OIOUR
#endifonly
#ifonly APPLE MANGO ORANGE CARROT
VALUE="JUICE" REQUIRED="yes" 
as df fg
#endifonly

【问题讨论】:

    标签: regex perl multiline


    【解决方案1】:
    use strict;
    use warnings;
    use 5.010;
    
    while (<DATA>) {
       my $rc = /#ifonly .+ SPINACH/ .. (my ($value) = /VALUE="([^"]*)"/);
       next unless $rc =~ /E0$/;
       say $value;
    }
    
    __DATA__
    #ifonly APPLE CARROT SPINACH
    VALUE="SALAD" REQUIRED="yes" 
    QW RETEWRT OIOUR
    #endifonly
    #ifonly APPLE MANGO ORANGE CARROT
    VALUE="JUICE" REQUIRED="yes" 
    as df fg
    #endifonly
    

    这使用了 brian d foy here 描述的一个小技巧。如链接所述,它使用标量 range operator / flipflop。

    【讨论】:

    • 另外,有点短:next unless (/#ifonly .+ SPINACH/ .. (my ($value) = /VALUE="([^"]*)"/)) =~ / E0$/; 但坦率地说,它破坏了我的缩进,所以我不会使用它。:) 那里也发生了很多事情,这可能不是最好的可维护性。
    • 非常酷的方法,并且(再一次)你发布的链接教会了我一些东西,所以谢谢你!
    • @canavanin 我有链接。他们都是!不客气——Effective Perler 是我最喜欢的 Perl 博客,所以很高兴在那里指导人们。
    【解决方案2】:

    如果你的文件很大(或者你想逐行阅读),你可以这样做:

    #!/usr/bin/perl
    
    use strict;
    use warnings;
    use Getopt::Long;
    
    my ($file, $keyword);
    
    # now get command line options (see Usage note below)
    GetOptions(
                "f=s" => \$file,
                "k=s" => \$keyword,
              );
    
    # if either the file or the keyword has not been provided, display a
    # help text and exit
    if (! $file || ! $keyword) {
       print STDERR<<EOF;
    
       Usage: script.pl -f filename -k keyword
    
    EOF
       exit(1);
    }
    
    my $found;         # indicator that the keyword has been found
    my $returned_word; # will store the word you want to retrieve
    
    open FILE, "<$file" or die "Cannot open file '$file': $!";
    while (<FILE>) {
       if (/$keyword/) {
          $found = 1;
       }
    
       # the following condition will be true between all lines that
       # start with '#ifonly' or '#endifonly' - but only if the keyword 
       # has been found!
       if (/^#ifonly/ .. /^#endifonly/ && $found) {
          if (/VALUE="(\w+)"/) { 
             $returned_word = $1;
             print "looking for $keyword --> found $returned_word\n";
    
             last; # if you want to get ALL values after the keyword
                   # remove the 'last' statement, as it makes the script
                   # exit the while loop
          }
       }
    }
    close FILE;
    

    【讨论】:

      【解决方案3】:

      您可以读取字符串中的文件内容,然后在字符串中搜索模式:

      my $file;    
      $file.=$_ while(<>);    
      if($file =~ /#ifonly.+?\bSPINACH\b.+?VALUE="(\w*)".+?#endifonly/s) {
              print $1;
      }
      

      您的原始正则表达式需要一些调整:

      • 您需要制作量词 不贪心。
      • 使用s 修饰符使. 也匹配换行符。

      Ideone Link

      【讨论】:

        【解决方案4】:

        这是基于触发器运算符的另一个答案:

        use strict;
        use warnings;
        use 5.010;
        
        while (<$file>)
        {
          if ( (/^#ifonly.*\bSPINACH\b/ .. /^#endifonly/) &&
               (my ($chosen) = /^VALUE="(\w+)"/) )
          {
            say $chosen;
          }
        }
        

        此解决方案将第二个测试应用于范围内的所有行。不需要 @Hugmeir 用来排除开始行和结束行的技巧,因为“内部”正则表达式 /^VALUE="(\w+)"/ 无论如何都无法匹配它们(我在所有正则表达式中添加了 ^ 锚点,以确保这一点) .

        【讨论】:

          【解决方案5】:

          两天前给出的一个答案中的这两行

          my $file;
          $file.=$_ while(<>);
          

          效率不高。 Perl 可能会以大块的形式读取文件,将这些块分成&lt;&gt; 的文本行,然后.= 会将这些行重新组合成一个大字符串。啜食文件会更有效。基本样式是更改输入记录分隔符\$。

          undef $/;
          $file = <>;
          

          模块File::Slurp;(参见perldoc File::Slurp)可能会更好。

          【讨论】:

            猜你喜欢
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            相关资源
            最近更新 更多