【问题标题】:Unique cumulative customers by each day每天的唯一累积客户
【发布时间】:2021-03-13 22:45:55
【问题描述】:

任务:获取每个拒绝原因和每天的唯一累积客户总数。

    Input data sample:
    +---------+--------------+------------+------+
    | Cust_Id |  Decline_Dt  |  Reason    | Days |
    +---------+--------------+------------+------+
    |   A     |  08-09-2020  |  Reason_1  |   0  |
    |   A     |  08-09-2020  |  Reason_1  |   1  | 
    |   A     |  08-09-2020  |  Reason_1  |   2  |
    |   A     |  08-09-2020  |  Reason_1  |   4  |
    |   B     |  08-09-2020  |  Reason_1  |   0  |
    |   B     |  08-09-2020  |  Reason_1  |   2  |
    |   B     |  08-09-2020  |  Reason_1  |   3  |
    |   C     |  08-09-2020  |  Reason_1  |   1  |
    +---------+--------------+------------+------+   
1) Decline_dt - The date on which the payment was declined. (Ignore it for this task)
2) Days - Indicates the # of days after the payment decline happened, the customer interacted with IVR channel. 
3) Reason - Indicates the payment decline reason 
    
    --Expected Output:
    +---------------+-----------+---------------+----------------------------+
    |   Reason      |   Days    | Unique_mtns   | total_cumulative_customers |
    +---------------+-----------+---------------+----------------------------+
    |   Reason_1    |   0       |   2           |           2                |              
    |   Reason_1    |   1       |   2           |           3                |
    |   Reason_1    |   2       |   2           |           3                | 
    |   Reason_1    |   3       |   1           |           3                | 
    |   Reason_1    |   4       |   1           |           3                | 
    +------------------------------------------------------------------------+

我的 Hive 查询:

select a.Reason
        , a.days
        -- , count(distinct a.cust_id) as unique_mtns
        , count(distinct a.cust_id) over (partition by Reason 
                                                order by a.days rows between unbounded preceding and current row) 
                                        as total_cumulative_customers
from table as a  
group by a.reason
        , a.days

输出(不正确):

+---------------+-----------+----------------------------+
|   Reason      |   Days    | total_cumulative_customers |
+---------------+-----------+----------------------------+
|   Reason_1    |   0       |               2            |              
|   Reason_1    |   1       |               2            |
|   Reason_1    |   2       |               2            | 
|   Reason_1    |   3       |               1            | 
|   Reason_1    |   4       |               1            | 
+--------------------------------------------------------+

理想情况下,我希望在没有 group by 的情况下执行窗口函数。 但是,我得到一个没有 group by 的错误。当我使用 group by 时,我没有得到累积的客户。

【问题讨论】:

    标签: sql hive count hiveql window-functions


    【解决方案1】:

    如果我没看错,您可以使用子查询来计算每个客户/原因元组的第一天,然后进行条件聚合:

    select reason, days, 
        count(distinct cust_id) as unique_mtns,
        sum(sum(case when days = min_days then 1 else 0 end)) 
            over(partition by reason order by days) as total_cumulative_customers
    from (
        select reason, cust_id, 
            min(days) over(partition by reason, cust_id) as min_days
        from mytable
    ) t
    group by reason, days
    

    【讨论】:

    • 某些列在外部查询中不可用。 min_days 也不存在于 group by 中,所以应该有另一个 sum 或者它应该抛出一个错误(也许在 Hive 中它的工作方式不同?)。这可以通过在内部查询中使用lag(0, 1, 1) over(partition by reason, cust_id order by days asc)标记第一天来解决。
    • 谢谢@GMB!在您通过添加另一个总和进行编辑后,您的解决方案可以正常工作。谢谢@astentx,通过添加另一个总和,我运行了查询。
    【解决方案2】:

    我建议使用row_number() 来枚举行或给定的客户和原因。您的代码在用户 ID 上使用了 count(distinct),这表明您可能在给定日期有重复项。

    这将是:

    select reason, days, count(distinct cust_id) as unique_mtns,
           sum(sum(case when seqnum = 1 then 1 else 0 end)) over (partition by reason order by days) as total_cumulative_customers
    from (select t.*,
                 row_number() over (partition by reason, cust_id order by days) as seqnum
          from t
         ) t
    group by reason, days
    order by reason, days;
    

    【讨论】:

    • 我在某一天没有重复。 count(*) 给我同样的结果。无论如何,您的解决方案也有效。谢谢!
    • @Ashish 。 . .您的问题使用 count(distinct) 给人的印象是您的数据有这样的重复。
    猜你喜欢
    • 2014-01-03
    • 2018-06-28
    • 1970-01-01
    • 2019-10-02
    • 1970-01-01
    • 2013-03-19
    • 1970-01-01
    • 1970-01-01
    • 2022-11-02
    相关资源
    最近更新 更多