【发布时间】:2017-06-13 21:48:52
【问题描述】:
在 Apache Pig(版本 0.16.x)中,通过数据集字段之一的现有值列表过滤数据集的最有效方法有哪些?
例如, (根据@inquisitive_mind 的提示更新)
输入:一个行分隔的文件,每行一个值 my_codes.txt
'110'
'100'
'000'
sample_data.txt
'110', 2
'110', 3
'001', 3
'000', 1
期望的输出
'110', 2
'110', 3
'000', 1
示例脚本
%default my_codes_file 'my_codes.txt'
%default sample_data_file 'sample_data.txt'
my_codes = LOAD '$my_codes_file' as (code:chararray)
sample_data = LOAD '$sample_data_file' as (code: chararray, point: float)
desired_data = FILTER sample_data BY code IN (my_codes.code);
错误:
Scalar has more than one row in the output. 1st : ('110'), 2nd :('100')
(common cause: "JOIN" then "FOREACH ... GENERATE foo.bar" should be "foo::bar" )
我也尝试过FILTER sample_data BY code IN my_codes;,但“IN”子句似乎需要括号。
我也试过FILTER sample_data BY code IN (my_codes);,但得到了错误:
一列需要从关系中投影出来才能用作标量
【问题讨论】:
-
跟进:如果现有列表与正在查询的数据集相比较小,则 Pig 中的 复制 连接比标准 JOIN 更有效。
标签: apache-pig