【问题标题】:COUNT() OVER conditioned on the CURRENT ROWCOUNT() OVER 以当前行为条件
【发布时间】:2017-04-04 03:53:58
【问题描述】:

给定代表一个任务的每一行,以及开始时间和结束时间,我如何使用窗口函数计算每个任务开始(包括其本身)时正在运行的任务(即开始和未结束)的数量COUNT OVER?窗口函数是正确的方法吗?

例如,给定表tasks:

task_id  start_time  end_time
   a         1          10
   b         2           5
   c         5          15
   d         8          13
   e        12          20
   f        21          30

计算running_tasks:

task_id  start_time  end_time  running_tasks
   a         1          10           1         # a
   b         2           5           2         # a,b
   c         5          15           2         # a,c (b has ended)
   d         8          13           3         # a,c,d
   e        12          20           3         # c,d,e (a has ended)
   f        21          30           1         # f (c,d,e have ended)

【问题讨论】:

    标签: sql google-bigquery window-functions


    【解决方案1】:
    select      task_id,start_time,end_time,running_tasks 
    
    from       (select      task_id,tm,op,start_time,end_time
    
                           ,sum(op) over 
                            (
                                order by    tm,op 
                                rows        unbounded preceding
                            ) as running_tasks 
    
                from       (select      task_id,start_time as tm,1 as op,start_time,end_time 
                            from        tasks 
    
                            union   all 
    
                            select      task_id,end_time as tm,-1 as op,start_time,end_time 
                            from        tasks 
                            ) t 
                )t 
    
    where       op = 1
    ;
    

    【讨论】:

    • 嘟嘟,谢谢-聪明的解决方案,它解决了具体的简化示例。我希望得到一个更通用的解决方案,其中窗口函数可以以当前行为条件 - 这可能吗?
    • @NewDev,我不明白你的意图,所以请提供一个例子。
    • 是否可以通过窗口函数COUNT 仅计算满足当前行条件的行。比如当前行的start_time = 8,窗口函数是否只能统计end_time > 8和start_time <=8的行?
    • 这个查询已经实现了一个等效的逻辑。每个点的运行任务数是包含当前行开始时间的行数。您还需要什么其他信息?
    • 我想我的意思是,不重复行。我会接受您的回答,看看是否可以更好地使用更复杂的示例来说明问题。
    【解决方案2】:

    您可以使用相关子查询,在这种情况下是自联接;不需要分析函数。启用standard SQL(取消选中 UI 中“显示选项”下的“使用旧版 SQL”)后,您可以运行此示例:

    WITH tasks AS (
      SELECT
        task_id,
        start_time,
        end_time
      FROM UNNEST(ARRAY<STRUCT<task_id STRING, start_time INT64, end_time INT64>>[
        ('a', 1, 10),
        ('b', 2, 5),
        ('c', 5, 15),
        ('d', 8, 13),
        ('e', 12, 20),
        ('f', 21, 30)
      ])
    )
    SELECT
      *,
      (SELECT COUNT(*) FROM tasks t2
       WHERE t.start_time >= t2.start_time AND
       t.start_time < t2.end_time) AS running_tasks
    FROM tasks t
    ORDER BY task_id;
    

    【讨论】:

    • “不需要分析函数”?这在您看来是一种优势?
    • 是的——向新的 SQL 用户解释分析函数通常比连接和聚合等概念更难。在这种情况下,OP 现在有答案,可以对问题给出两种不同的观点,这很好:)
    • 知道了,但我至少要加一个小标志——“警告!前方有性能危险!”
    • 谢谢 Elliott...我知道这种方法,这就是为什么这个问题是针对窗口函数和分析的
    【解决方案3】:

    正如 Elliott 所提到的 - “向新用户解释分析函数通常更困难”,即使是老用户也并不总是 100% 擅长它(虽然非常接近它)!
    所以,虽然 Dudu Markovitz 的回答很棒——不幸的是,它仍然是不正确的(至少根据我对问题的理解)。不正确的情况是您在同一个 start_time 启动了多个任务 - 所以这些任务的“运行任务”结果错误

    作为一个例子 - 考虑下面的例子:

    task_id  start_time  end_time
       a         1          10
       aa        1           2
       aaa       1           8
       b         2           5
       c         5          15
       d         8          13
       e        12          20
       f        21          30
    

    我想,你会期待以下结果:

    task_id  start_time  end_time  running_tasks
       a         1          10           3         # a,aa,aaa
       aa        1           2           3         # a,aa,aaa
       aaa       1           8           3         # a,aa,aaa
       b         2           5           3         # a,aaa,b (aa has ended)
       c         5          15           3         # a,aaa,c (b has ended)
       d         8          13           3         # a,c,d (aaa has ended)
       e        12          20           3         # c,d,e (a has ended)
       f        21          30           1         # f (c,d,e have ended)     
    

    如果你用 Dudu 的代码试试 - 你会得到下面的代替

    task_id  start_time  end_time  running_tasks
       a         1          10           1        
       aa        1           2           2        
       aaa       1           8           3        
       b         2           5           3        
       c         5          15           3        
       d         8          13           3        
       e        12          20           3        
       f        21          30           1        
    

    您可以看到任务 a 和 aa 的结果错误。
    原因是因为使用了ROWS UNBOUNDED PRECEDING 而不是RANGE UNBOUNDED PRECEDING - 细微但非常重要的细微差别!

    所以下面的查询会给你正确的结果

    SELECT  task_id,start_time,end_time,running_tasks 
    FROM  (
      SELECT  
        task_id, tm, op, start_time, end_time,
        SUM(op) OVER (ORDER BY  tm ,op RANGE UNBOUNDED PRECEDING) AS running_tasks 
      FROM  (
        SELECT  
          task_id, start_time AS tm, 1 AS op, start_time, end_time 
        FROM  tasks UNION  ALL 
        SELECT  
          task_id, end_time AS tm, -1 AS op, start_time, end_time 
        FROM  tasks 
      ) t 
    )t 
    WHERE  op = 1
    ORDER BY start_time       
    

    快速总结:
    ROWS UNBOUNDED PRECEDING - 根据行的位置设置窗口框架
    而
    RANGE UNBOUNDED PRECEDING - 根据行值设置窗口框架

    再次 - 正如 Elliott 所提到的 - 这比 JOIN 概念要复杂得多 - 但它值得(因为它比连接更有效) - 了解更多关于 Window Frame Clause 和 ROWS 与 RANGE 的使用

    【讨论】:

      猜你喜欢
      • 2019-02-08
      • 1970-01-01
      • 1970-01-01
      • 2021-05-06
      • 2022-12-02
      • 2019-10-22
      • 1970-01-01
      • 1970-01-01
      • 2021-12-15
      相关资源
      最近更新 更多