【问题标题】:multi file insert from hive table not working?从蜂巢表插入多文件不起作用?
【发布时间】:2016-11-25 16:49:30
【问题描述】:

您好,我在 HBase 上的一个配置单元表中有 200 GB 的数据。 我必须从当前仅尝试 3 个文件的表中创建 142 个不同的文件。

我想同时运行所有查询以并行运行。 我正在尝试从 hive 表中插入多文件但得到解析异常。

这是我正在尝试的查询。

FROM  hbase_table_FinancialLineItem

INSERT OVERWRITE LOCAL DIRECTORY '/hadoop/user/m6034690/FSDI/FinancialLineItem/Japan.txt'
ROW FORMAT DELIMITED
FIELDS TERMINATED BY '\t'
STORED AS TEXTFILE
select * from hbase_table_FinancialLineItem WHERE FilePartition='Japan'

INSERT OVERWRITE LOCAL DIRECTORY '/hadoop/user/m6034690/FSDI/FinancialLineItem/SelfSourcedPrivate.txt'
ROW FORMAT DELIMITED
FIELDS TERMINATED BY '\t'
STORED AS TEXTFILE
select * from hbase_table_FinancialLineItem WHERE FilePartition='SelfSourcedPrivate'


INSERT OVERWRITE LOCAL DIRECTORY '/hadoop/user/m6034690/FSDI/FinancialLineItem/ThirdPartyPrivate.txt'
ROW FORMAT DELIMITED
FIELDS TERMINATED BY '\t'
STORED AS TEXTFILE
select * from hbase_table_FinancialLineItem WHERE FilePartition='ThirdPartyPrivate';

运行后我遇到了错误。

FAILED: ParseException line 7:9 missing EOF at 'from' near '*'

【问题讨论】:

  • ParseException 是因为您在每个 INSERT 末尾缺少分号
  • 我想导入本地目录。如果我提供半列,那么每个查询将独立运行,不像多表插入。
  • 我不认为它会那样工作。如果要并行运行所有 hive insert,请尝试将其作为 oozie hive action 作业运行
  • 是的,我也打算这样做。另外,我无法在 hive 中创建分区,因为它指向 HBase。只是一个问题,知道从那里获取所有 200 gb 数据需要多长时间oozie 操作,如果我打算创建 142 个文件?
  • 取决于可用资源。有多少内核和内存 -> 检查 yarn web 控制台

标签: hadoop hive


【解决方案1】:

我认为在每次插入覆盖的末尾添加此FROM hbase_table_FinancialLineItem; 即可解决。

【讨论】:

    猜你喜欢
    • 2018-12-21
    • 2019-02-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多