【问题标题】:Regex to parse html for sentences?正则表达式来解析 html 的句子?
【发布时间】:2013-05-22 06:08:43
【问题描述】:

我知道 HTML:Parser 是一个东西,通过阅读,我意识到尝试使用正则表达式解析 html 通常是一种次优的做事方式,但是对于 Perl 类,我目前正在尝试使用常规表达式(希望只是一个匹配)来识别和存储保存的 html 文档中的句子。最终,我希望能够计算出句子的数量、单词/句子以及页面上单词的平均长度。

目前,我只是尝试隔离“>”之后和“。”之前的内容,只是为了看看它隔离了什么,但我无法让代码运行,即使在操作常规时表达。所以我不确定问题是在正则表达式中,还是在其他地方,或者两者兼而有之。任何帮助将不胜感激!

#!/usr/bin/perl
#new
use CGI qw(:standard);
print header;

open FILE, "< sample.html ";
$html = join('', <FILE>);
close FILE;

print "<pre>";

###Main Program###
&sentences;

###sentence identifier sub###

sub sentences {
@sentences;
while ($html =~ />[^<]\. /gis) {
    push @sentences, $1;
}
#for debugging, comment out when running    
    print join("\n",@sentences);
}

print "</pre>";

【问题讨论】:

  • 您能说说您遇到了什么错误吗?无法运行代码是什么意思?
  • 我希望我能,我收到的唯一错误是服务器错误 (500),服务器为我提供了从缺少
     语句到不正确语法到缺少括号的所有内容
  • 你仍然从不使用 Regex 解析 HTML。

标签: html regex perl parsing


【解决方案1】:

你的正则表达式应该是/&gt;[^&lt;]*?./gis

*? 表示匹配零个或多个非贪婪。就目前而言,您的正则表达式将仅匹配一个非

可能还有其他问题。

现在阅读this

【讨论】:

    【解决方案2】:

    第一个改进是写$html =~ /&gt;([^&lt;.]+)\. /gs,你需要捕捉与父母的匹配,并允许每个句子超过1个字母;--)

    这并没有得到所有的句子,只是每个元素中的第一个。

    更好的方法是捕获所有文本,然后从每个片段中提取句子

    while( $html=~ m{>([^<]*<}g) { push @text_content, $1}; 
    foreach (@text_content) { while( m{([^.]*)\.}gs) { push @sentences, $1; } }
    

    (未经测试,因为现在是清晨,咖啡在召唤)

    关于使用正则表达式解析 HTML 的所有常见警告都适用,最明显的是文本中存在“>”。

    【讨论】:

    • 感谢您的解释,我现在更清楚地了解正则表达式的形成方式了!
    【解决方案3】:

    我认为这或多或少可以满足您的需求。请记住,此脚本仅查看 p 标签内的文本。文件名作为命令行参数 (shift) 传入。

    #!/usr/bin/perl
    
     use strict;
     use warnings;
     use HTML::Grabber;
    
     my $file_location = shift;
     print "\n\nfile: $file_location";
     my $totalWordCount = 0;
     my $sentenceCount = 0;
     my $wordsInSentenceCount = 0;
     my $averageWordsPerSentence = 0;
     my $char_count = 0;
     my $contents;
     my $rounded;
     my $rounded2;
    
     open ( my $file, '<', $file_location  ) or die "cannot open < file: $!";
    
        while( my $line = <$file>){
              $contents .= $line;
      }      
     close( $file );
     my $dom = HTML::Grabber->new( html => $contents );
    
     $dom->find('p')->each( sub{
        my $p_tag = $_->text;
    
        ++$totalWordCount while $p_tag =~ /\S+/g;
    
    
        while ($p_tag =~ /[.!?]+/g){
                  $p_tag =~ s/\s//g;
                  $char_count += (length($p_tag));
                  $sentenceCount++;  
              }
         });     
    
    
               print "\n Total Words: $totalWordCount\n";
               print " Total Sentences: $sentenceCount\n";
               $rounded = $totalWordCount / $sentenceCount;
               print  " Average words per sentence: $rounded.\n\n";
               print " Total Characters: $char_count.\n";
               my $averageCharsPerWord = $char_count / $totalWordCount  ;
    
               $rounded2 = sprintf("%.2f", $averageCharsPerWord );
    
               print  " Average words per sentence: $rounded2.\n\n";
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2021-07-20
      • 2014-12-06
      • 2012-09-12
      • 2014-05-16
      • 1970-01-01
      • 2014-06-08
      相关资源
      最近更新 更多