【问题标题】:add a new column for unique ID in hive table在 hive 表中为唯一 ID 添加一个新列
【发布时间】:2016-09-01 10:46:59
【问题描述】:

我在 hive 中有一个表,其中包含两列:session_idduration_time,如下所示:

|| session_id || duration||

    1               14          
    1               10      
    1               20          
    1               10          
    1               12          
    1               16          
    1               8       
    2               9           
    2               6           
    2               30          
    2               22

我想在以下情况下添加一个具有唯一 ID 的新列:

session_id 正在改变duration_time > 15

我希望输出是这样的:

session_id      duration    unique_id
1               14          1
1               10          1
1               20          2
1               10          2
1               12          2
1               16          3
1               8           3
2               9           4
2               6           4
2               30          5
2               22          6

任何想法如何在 hive QL 中做到这一点?

谢谢!

【问题讨论】:

    标签: hadoop hive unique calculated-columns unique-key


    【解决方案1】:

    SQL 表代表无序 集合。您需要一列指定值的顺序,因为您似乎关心顺序。例如,这可以是 id 列或 created-at 列。

    您可以使用累积和来做到这一点:

    select t.*,
           sum(case when duration > 15 or seqnum = 1 then 1 else 0 end) over
               (order by ??) as unique_id
    from (select t.*,
                 row_number() over (partition by session_id order by ??) as seqnum
          from t
         ) t;
    

    【讨论】:

    • 它与 ID 列完美配合,我已经在 impala 上运行它。非常感谢
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-09-03
    • 1970-01-01
    • 1970-01-01
    • 2017-01-01
    • 1970-01-01
    相关资源
    最近更新 更多