【问题标题】:How to generate a custom schema from a relation in Pig?如何从 Pig 中的关系生成自定义模式?
【发布时间】:2011-04-15 19:59:51
【问题描述】:

我有一个描述各种文章中单词的 tf-idf 值的模式。 它的描述如下:

tfidf_relation: {word: chararray,id: bytearray,tfidf: double}

以下是此类数据的示例:

(cat,article_one,0.13515503603605478)
(cat,article_two,0.4054651081081644)
(dog,article_one,0.3662040962227032)
(apple,article_three,0.3662040962227032)
(orange,article_three,0.3662040962227032)
(parrot,article_one,0.13515503603605478)
(parrot,article_three,0.13515503603605478)

我想以一种形式获得输出: 猫 article_one 0.13515503603605478,article_two 0.4054651081081644 等等。 问题是,我如何从中建立一个包含单词字段和 id 和 tfidf 字段元组的关系? 像这样的:

X = FOREACH tfidf_relation GENERATE word, (id, tfidf);

不起作用。正确的语法是什么?

【问题讨论】:

    标签: hadoop mapreduce apache-pig


    【解决方案1】:

    试试这个:

        t = LOAD 'input/file' USING PigStorage(',') as (word: chararray,id: bytearray,tfidf: double);
        u = group t by word;
        dump u;
    

    输出将是

        (cat,{(cat,article_two,0.4054651081081644),(cat,article_one,0.13515503603605478)})
        (dog,{(dog,article_one,0.3662040962227032)})
        (apple,{(apple,article_three,0.3662040962227032)})
        (orange,{(orange,article_three,0.366204096222703)})
        (parrot,{(parrot,article_three,0.13515503603605478),
        (parrot,article_one,0.13515503603605478)})
    

    我希望这就是你要找的。​​p>

    【讨论】:

      【解决方案2】:
      X = FOREACH tfidf_relation GENERATE word, {(id, tfidf)};
      

      这可能是你需要的。

      【讨论】:

      • Wojtek,我尝试了不同形式的解决方案 - 每次在符号“{”处出现解析错误。
      • 好的,我想我已经使用 Java 例程嵌入实现了这一点。但是正确的 Pig 语法(以及可能性本身)仍然很有趣。
      • X.(id, tfidf) 怎么样?您总是可以按单词分组然后展平分组(如果我记得的话,就是这样),但可能更容易编写快速 UDF。
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2023-04-01
      • 1970-01-01
      • 2011-11-18
      • 1970-01-01
      • 1970-01-01
      • 2012-06-01
      • 1970-01-01
      相关资源
      最近更新 更多