【问题标题】:How to read ORC file in hadoop streaming?如何在 hadoop 流中读取 ORC 文件?
【发布时间】:2015-11-25 07:06:43
【问题描述】:

我想在 Python 上的 mapreduce 中读取 ORC 文件。我尝试运行它:

hadoop jar /usr/lib/hadoop/lib/hadoop-streaming-2.6.0.2.2.6.0-2800.jar 
-file /hdfs/price/mymapper.py 
-mapper '/usr/local/anaconda/bin/python mymapper.py' 
-file /hdfs/price/myreducer.py 
-reducer '/usr/local/anaconda/bin/python myreducer.py' 
-input /user/hive/orcfiles/* 
-libjars /usr/hdp/2.2.6.0-2800/hive/lib/hive-exec.jar 
-inputformat org.apache.hadoop.hive.ql.io.orc.OrcInputFormat 
-numReduceTasks 1 
-output /user/hive/output

但我得到错误:

-inputformat : class not found : org.apache.hadoop.hive.ql.io.orc.OrcInputFormat

我发现了一个类似的问题OrcNewInputformat as a inputformat for hadoop streaming,但答案不清楚

请举例说明如何在 hadoop 流中正确读取 ORC 文件。

【问题讨论】:

    标签: python hadoop streaming orc


    【解决方案1】:

    这是我使用 ORC 分区 Hive 表作为输入的示例之一:

        hadoop jar /usr/hdp/2.2.4.12-1/hadoop-mapreduce/hadoop-streaming-2.6.0.2.2.4.12-1.jar \
    -libjars /usr/hdp/current/hive-client/lib/hive-exec.jar \
    -Dmapreduce.task.timeout=0 -Dmapred.reduce.tasks=1 \
    -Dmapreduce.job.queuename=default \
     -file RStreamMapper.R RStreamReducer2.R \
    -mapper "Rscript RStreamMapper.R" -reducer "Rscript RStreamReducer2.R" \
    -input /hive/warehouse/asv.db/rtd_430304_fnl2 \
    -output /user/Abhi/MRExample/Output \
    -inputformat org.apache.hadoop.hive.ql.io.orc.OrcInputFormat 
    -outputformat org.apache.hadoop.hive.ql.io.orc.OrcOutputFormat
    

    这里/apps/hive/warehouse/asv.db/rtd_430304_fnl2是HIVE表后台ORC数据存放位置的路径。休息一下,我需要为流式传输和 HIVE 提供适当的 jar。

    【讨论】:

      猜你喜欢
      • 2016-07-28
      • 2015-12-19
      • 1970-01-01
      • 1970-01-01
      • 2017-08-07
      • 2016-12-28
      • 2017-08-14
      • 2015-07-17
      • 2018-10-26
      相关资源
      最近更新 更多