【问题标题】:How can I get the sum(value) on the latest gather_time per group(name,col1) in PostgreSQL?如何在 PostgreSQL 中获取每个组(名称,col1)的最新收集时间的总和(值)?
【发布时间】:2016-04-07 19:08:59
【问题描述】:

实际上,我在下面的线程中得到了关于类似问题的一个很好的答案,但我需要针对不同数据集的更多解决方案。

How to get the latest 2 rows ( PostgreSQL )

数据集有历史数据,我只想在最新的gather_time获取组的总和(值)。 最终结果应该如下:

 name  | col1 |     gather_time     | sum
-------+------+---------------------+-----
 first | 100  | 2016-01-01 23:12:49 |   6
 first | 200  | 2016-01-01 23:11:13 |   4

但是,我只能看到一组(first-100)的数据,下面的查询意味着第二组(first-200)没有数据。 问题是我需要每组获得一排。 组的数量可以变化。

select name,col1,gather_time,sum(value) 
from testtable
group by name,col1,gather_time
order by gather_time desc
limit 2;

 name  | col1 |     gather_time     | sum
-------+------+---------------------+-----
 first | 100  | 2016-01-01 23:12:49 |   6
 first | 100  | 2016-01-01 23:11:19 |   6
(2 rows)

你能建议我完成这个要求吗?

数据集

create table testtable
(
name varchar(30),
col1 varchar(30),
col2 varchar(30),
gather_time timestamp,
value integer
);


insert into testtable values('first','100','q1','2016-01-01 23:11:19',2);
insert into testtable values('first','100','q2','2016-01-01 23:11:19',2);
insert into testtable values('first','100','q3','2016-01-01 23:11:19',2);
insert into testtable values('first','200','t1','2016-01-01 23:11:13',2);
insert into testtable values('first','200','t2','2016-01-01 23:11:13',2);
insert into testtable values('first','100','q1','2016-01-01 23:11:11',2);
insert into testtable values('first','100','q1','2016-01-01 23:12:49',2);
insert into testtable values('first','100','q2','2016-01-01 23:12:49',2);
insert into testtable values('first','100','q3','2016-01-01 23:12:49',2);

select * 
from testtable 
order by name,col1,gather_time;

 name  | col1 | col2 |     gather_time     | value
-------+------+------+---------------------+-------
 first | 100  | q1   | 2016-01-01 23:11:11 |     2
 first | 100  | q2   | 2016-01-01 23:11:19 |     2
 first | 100  | q3   | 2016-01-01 23:11:19 |     2
 first | 100  | q1   | 2016-01-01 23:11:19 |     2
 first | 100  | q3   | 2016-01-01 23:12:49 |     2
 first | 100  | q1   | 2016-01-01 23:12:49 |     2
 first | 100  | q2   | 2016-01-01 23:12:49 |     2
 first | 200  | t2   | 2016-01-01 23:11:13 |     2
 first | 200  | t1   | 2016-01-01 23:11:13 |     2

【问题讨论】:

    标签: postgresql


    【解决方案1】:

    一种选择是将您的原始表连接到一个表,该表仅包含每个name、col1 组的最新gather_time 记录。然后你可以对每个组取value 列的总和,得到你想要的结果集。

    SELECT t1.name, t1.col1, MAX(t1.gather_time) AS gather_time, SUM(t1.value) AS sum
    FROM testtable t1 INNER JOIN
    (
        SELECT name, col1, col2, MAX(gather_time) AS maxTime
        FROM testtable
        GROUP BY name, col1, col2
    ) t2
    ON t1.name = t2.name AND t1.col1 = t2.col1 AND t1.col2 = t2.col2 AND
        t1.gather_time = t2.maxTime
    GROUP BY t1.name, t1.col1
    

    如果您想在 WHERE 子句中使用子查询,就像您在 OP 中尝试的那样,以限制仅包含最新 gather_time 的记录,那么您可以尝试以下操作:

    SELECT name, col1, gather_time, SUM(value) AS sum
    FROM testtable t1
    WHERE gather_time =
    (
        SELECT MAX(gather_time) 
        FROM testtable t2
        WHERE t1.name = t2.name AND t1.col1 = t2.col1
    )
    GROUP BY name, col1
    

    【讨论】:

    • 为什么专家解决方案看起来很简单......非常好。太感谢了。您认为我提出的问题等同于您的解决方案吗?
    • 我个人更喜欢join解决方案,无论是从美学角度还是性能角度。
    • gather_time 列应该放在 Group by 我猜。好的。非常感谢!
    猜你喜欢
    • 2021-11-15
    • 1970-01-01
    • 2015-12-01
    • 2020-01-17
    • 1970-01-01
    • 2021-11-04
    • 2021-08-10
    • 1970-01-01
    • 2016-11-23
    相关资源
    最近更新 更多