【问题标题】:Perl - find and save in an associative array word and word contextPerl - 在关联数组单词和单词上下文中查找并保存
【发布时间】:2014-05-21 10:17:51
【问题描述】:

我有一个这样的数组(这只是一个小概述,但它有 2000 多行这样的行):

@list = (
        "affaire,chose,question",
        "cause,chose,matière",
);

我想要这个输出:

%te = (
affaire => "chose", "question",
chose => "affaire", "question", "cause", "matière", 
question => "affaire", "chose",
cause => "chose", "matière",
matière => "cause", "chose"
);

我已经创建了这个脚本,但效果不是很好,而且我认为它太复杂了..

use Data::Dumper;
@list = (
        "affaire,chose,question",
        "cause,chose,matière",
);

%te;

for ($a = 0; $a < @list; $a++){
    @split_list = split (/,/,$list[$a]);
}

foreach $elt (@split_list){
print "SPLIT ELT : $split_list[$elt]\n";

for ($i = 0; $i < @list; $i++){

    $test = $list[$i]; #$test = "affaire,chose,question"

    if (exists $te{$split_list[$elt]}){ #if exists affaire in %te

        @t = split (/,/,$test); # @t = affaire chose question
        print "T : @t\n";

        @temp = grep(!/$split_list[$elt]/, @t); 
        print "GREP : @temp\n";#@temp = chose question

        @fin = join(', ', @temp); #@fin = chose, question;

        for ($k = 0; $k < @fin; $k++){
            $te{$split_list[$elt]} .= $fin[$k]; #affaire => chose, question
        }

    }
    else {

                @t = split (/,/,$test); # @t = affaire chose question
        print "T : @t\n";

        @temp = grep(!/$split_list[$elt]/, @t); 
        print "GREP : @temp\n";#@temp = chose question

        @fin = join(', ', @temp); #@fin = chose, question;

        for ($k = 0; $k < @fin; $k++){
                $te{$split_list[$elt]} = $fin[$k];
                }
    }
}

}



print Dumper \%te;

输出:

SPLIT ELT : cause
T : affaire chose question
GREP : affaire chose question
T : cause chose matière
GREP : chose matière
SPLIT ELT : cause
T : affaire chose question
GREP : affaire chose question
T : cause chose matière
GREP : chose matière
SPLIT ELT : cause
T : affaire chose question
GREP : affaire chose question
T : cause chose matière
GREP : chose matière
$VAR1 = {
          'cause' => 'affaire, chose, questionchose, matièreaffaire, chose, questionchose, matièreaffaire, chose, questionchose, matière'
        };

【问题讨论】:

    标签: arrays perl nlp associative-array


    【解决方案1】:

    对于@list中的每个元素,将其拆分为,,并将每个字段用作%te的键,将其他字段推送到该键的值:

    #!/usr/bin/perl
    
    use strict;
    use warnings;
    
    use Data::Dumper;
    
    my @list = (
        "affaire,chose,question",
        "cause,chose,matière",
    );
    
    my %te;
    
    foreach my $str (@list) {
        my @field = split /,/, $str;
        foreach my $key (@field) {
            my @other = grep { $_ ne $key } @field;
            push @{$te{$key}}, @other;
        }
    }
    
    print Dumper(\%te);
    

    输出:

    $ perl t.pl
    $VAR1 = {
              'question' => [
                              'affaire',
                              'chose'
                            ],
              'affaire' => [
                             'chose',
                             'question'
                           ],
              'matière' => [
                              'cause',
                              'chose'
                            ],
              'cause' => [
                           'chose',
                           'matière'
                         ],
              'chose' => [
                           'affaire',
                           'question',
                           'cause',
                           'matière'
                         ]
            };
    

    【讨论】:

    • 很好... +1 为您的所有解决方案 - 我检查了编辑历史:-) 检查像这样的 HoA 中键的特定值的最佳实践或模式是什么? List::Util 和普通哈希访问的某种组合?
    • @G.Cito 我猜List::Util 和List::MoreUtils 在这种情况下都很有用。
    • 或者最近发现(至少是我)的奇妙List::AllUtils!不再有use List::Utils oops use List::Util oops use List::MoreUtils oops use List::Util ; use List::MoreUtils ... :-),干杯
    【解决方案2】:

    我认为我明白你在做什么:索引单词之间的语义链接,然后是同义词列表。我对么? :-)

    如果一个词出现在多个同义词列表中,那么为该词创建一个哈希条目,该词作为键并使用它最初是同义词的关键字作为值......或类似的东西。使用数组的散列 - 如@Lee Duhem 的解决方案 - 你会得到每个关键词的同义词列表(数组)。这是一种常见的模式。不过,你最终会得到很多哈希条目。

    我一直在使用@miygawa 的一个名为Hash::MultiValue 的简洁模块,它采用不同的方法来访问与每个散列键关联的值列表:多值散列。一些不错的功能是您可以从多值散列动态创建数组引用的散列,“展平”散列,编写回调以使用-&gt;each() 方法,以及其他简洁的东西,因此它非常灵活。我相信该模块没有依赖关系(除了测试)。另外它是由@miyagawa(和其他贡献者)提供的,所以使用它并阅读它对你有好处:-)

    我不是专家,我不确定它是否适合您想要的 - 作为 Lee 方法的变体,您可能会有类似的内容:

    #!/usr/bin/env perl
    use strict;
    use warnings;
    use Hash::MultiValue;
    
    my $words_hash = Hash::MultiValue->new();
    
    # set up the mvalue hash
    for my $words (<DATA>) {
      my @synonyms = split (',' , $words) ; 
      $words_hash->add( shift @synonyms => (@synonyms[0..$#synonyms]) ) ;
    };
    
    for my $key (keys %{ $words_hash } ) {
      print "$key --> ", join(", ",  $words_hash->get_all($key)) ;
    };
    
    print "\n";
    
    sub synonmize {
      my $bonmot = shift;
      my @bonmot_syns ;
    
      # check key "$bonmot" for word to search and show values
      push @bonmot_syns , $words_hash->get_all($bonmot);
    
      # now grab values but leave out synonym's synonyms
      foreach (keys %{ $words_hash } ) {
        if ($_ !~ /$bonmot/ && grep {/$bonmot/} $words_hash->get_all($_)) {
          push @bonmot_syns, grep {!/$bonmot/} $words_hash->get_all($_);
        }
      }
    
      # show the keys with values containing target word
      $words_hash->each(
        sub { push @bonmot_syns,  $_[0] if grep /$bonmot/ ,  @_[1..$#_] ; }
      );
    
      chomp @bonmot_syns ;
      print "synonymes pour \"$bonmot\": @bonmot_syns \n" ;
    }
    
    # find synonyms 
    synonmize("chose");
    synonmize("truc");
    synonmize("matière");
    
    __DATA__
    affaire,chose,question
    cause,chose,matière
    chose,truc,bidule
    fille,demoiselle,femme,dame
    

    输出:

    fille --> demoiselle, femme, dame
    affaire --> chose, question
    cause --> chose, matière
    chose --> truc, bidule
    
    synonymes pour "chose": truc bidule question matière affaire cause 
    synonymes pour "truc": bidule chose 
    synonymes pour "matière": chose cause
    

    Tie::Hash::MultiValue 是另一种选择。感谢@Lee 提供快速干净的解决方案:-)

    【讨论】:

    • argh ...再次阅读此代码我发现这是一种非常迂回的方法-如果更多人像我一样滥用事物,则可能是反模式... :-\ 因为Plack 取决于在这个模块上,可以看到正确的用法。您可以从 CPAN 中学到很多东西!
    • 它很有用!我不知道..太棒了!
    猜你喜欢
    • 2016-07-08
    • 2012-07-07
    • 1970-01-01
    • 2015-12-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-10-31
    • 2021-12-14
    相关资源
    最近更新 更多