【问题标题】:Hive - external (dynamically) partitioned tableHive - 外部(动态)分区表
【发布时间】:2013-07-26 13:06:02
【问题描述】:

我在 MySQL 中有一个表。 nas_comps。

select comp_code, count(leg_id) from nas_comps_01012011_31012011 n group by comp_code;
comp_code     count(leg_id)
'J'           20640
'Y'           39680

首先,我使用 Sqoop 将数据导入到 HDFSHadoop 版本 1.0.2):

sqoop import --connect jdbc:mysql://172.25.37.135/pros_olap2 \
--username hadoopranch \
--password hadoopranch \
--query "select * from nas_comps where dep_date between '2011-01-01' and '2011-01-10' AND \$CONDITIONS" \
-m 1 \
--target-dir /pros/olap2/dataimports/nas_comps

然后,我创建了一个外部的分区 Hive 表:

/*shows the partitions on 'describe' but not 'show partitions'*/
create external table  nas_comps(DS_NAME string,DEP_DATE string,
                                 CRR_CODE string,FLIGHT_NO string,ORGN string,
                                 DSTN string,PHYSICAL_CAP int,ADJUSTED_CAP int,
                                 CLOSED_CAP int)
PARTITIONED BY (LEG_ID int, month INT, COMP_CODE string)
location '/pros/olap2/dataimports/nas_comps'

分区列在描述时显示:

hive> describe extended nas_comps;
OK
ds_name string
dep_date        string
crr_code        string
flight_no       string
orgn    string
dstn    string
physical_cap    int
adjusted_cap    int
closed_cap      int
leg_id  int
month   int
comp_code       string

Detailed Table Information      Table(tableName:nas_comps, dbName:pros_olap2_optim, 
owner:hadoopranch, createTime:1374849456, lastAccessTime:0, retention:0, 
sd:StorageDescriptor(cols:[FieldSchema(name:ds_name, type:string, comment:null), 
FieldSchema(name:dep_date, type:string, comment:null), FieldSchema(name:crr_code, 
type:string, comment:null), FieldSchema(name:flight_no, type:string, comment:null), 
FieldSchema(name:orgn, type:string, comment:null), FieldSchema(name:dstn, type:string, 
comment:null), FieldSchema(name:physical_cap, type:int, comment:null), 
FieldSchema(name:adjusted_cap, type:int, comment:null), FieldSchema(name:closed_cap, 
type:int, comment:null), FieldSchema(name:leg_id, type:int, comment:null), 
FieldSchema(name:month, type:int, comment:null), FieldSchema(name:comp_code, type:string, 
comment:null)], location:hdfs://172.25.37.21:54300/pros/olap2/dataimports/nas_comps, 
inputFormat:org.apache.hadoop.mapred.TextInputFormat, 
outputFormat:org.apache.hadoop.hive.ql.io.HiveIgnoreKeyTextOutputFormat, compressed:false, 
numBuckets:-1, serdeInfo:SerDeInfo(name:null, 
serializationLib:org.apache.hadoop.hive.serde2.lazy.LazySimpleSerDe, parameters:
{serialization.format=1}), bucketCols:[], sortCols:[], parameters:{}), partitionKeys:
[FieldSchema(name:leg_id, type:int, comment:null), FieldSchema(name:month, type:int,
comment:null), FieldSchema(name:comp_code, type:string, comment:null)], 
parameters:{EXTERNAL=TRUE, transient_lastDdlTime=1374849456}, viewOriginalText:null, 
viewExpandedText:null, tableType:EXTERNAL_TABLE)

但我不确定是否创建了分区,因为:

hive> show partitions nas_comps;
OK
Time taken: 0.599 seconds


select count(1) from nas_comps;

返回 0 条记录

如何创建具有动态分区的外部 Hive 表?

【问题讨论】:

    标签: hive hiveql


    【解决方案1】:

    动态分区

    在将记录插入配置单元表期间动态添加分区。

    1. 仅支持插入语句。
    2. load data 语句不支持。
    3. 在将数据插入 hive 表之前需要启用动态分区设置。 hive.exec.dynamic.partition.mode=nonstrict 默认值为strict hive.exec.dynamic.partition=true 默认值为false

    动态分区查询

    SET hive.exec.dynamic.partition.mode=nonstrict;
    SET hive.exec.dynamic.partition=true;
    INSERT INTO table_name PARTITION (loaded_date)
    select * from table_name1 where loaded_date = 20151217
    

    这里loaded_date = 20151217是分区和它的值。

    限制:

    1. 动态分区仅适用于上述语句。
    2. 它将根据从table_name1loaded_date 列中选择的数据动态创建分区;

    如果您的条件不符合上述条件,那么:

    首先创建一个分区表,然后这样做:

    ALTER TABLE table_name ADD PARTITION (DS_NAME='partname1',DATE='partname2'); 
    

    或请使用此Link 创建动态分区。

    【讨论】:

    • 是的,我已经检查过了,但这些不是动态分区——仍然需要为分区提供值。
    • 对,通过shell脚本运行。你可以在shell脚本中为分区创建一个变量,并在alter table命令中传递,否则目前没有可用的选项:(
    【解决方案2】:

    Hive 不会以这种方式为您创建分区。
    只需创建一个按所需分区键分区的表,然后从外部表执行insert overwrite table 到新的分区表(设置hive.exec.dynamic.partition=truehive.exec.dynamic.partition.mode=nonstrict)。

    如果您必须保持表在外部分区,则必须手动创建目录(每个分区 1 个目录,名称应为 PARTION_KEY=VALUE) 然后使用MSCK REPAIR TABLE table_name;command

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2023-03-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-10-01
      • 1970-01-01
      相关资源
      最近更新 更多