【发布时间】:2022-01-22 21:41:01
【问题描述】:
我正在对 postgres 13 + postgis 3 + timescaleDB 2 数据库中的一些 timeseris GPS 数据进行复杂聚合。我正在查看的表每天有几百万个条目,我想在几个月内进行聚合(每天一行,每个 gps_id,每个组间隙 ID)。
假设我创建了一个函数来执行聚合:
--pseudo code, won't actually work...
CREATE FUNCTION my_agg_func(starttime, endtime)
AS
WITH gps_window AS
(SELECT gps.id,
gps.geom,
gps.time,
-- find where there are 1 hour gaps in data
lag(ais.time) OVER (PARTITION BY gps.id ORDER BY gps.time) <= (gps.time - '01:00:00'::interval) AS time_step,
-- find where there are 0.1 deg gaps in position
st_distance(gps.geom, lag(gps.geom) OVER (PARTITION BY gps.id ORDER BY gps.time)) >= 0.1 AS dist_step
FROM gps
WHERE gps.time BETWEEN starttime AND endtime
), groups AS (
SELECT gps_window.id,
gps_window.geom,
gps_window.time,
count(*) FILTER (WHERE gps_window.time_step) OVER (PARTITION BY gps_window.id ORDER BY gps_window.time) AS time_grp,
count(*) FILTER (WHERE gps_window.dist_step) OVER (PARTITION BY gps_window.id ORDER BY gps_window.time) AS dist_grp
FROM gps_window
--get rid of duplicate points
WHERE gps_window.dist > 0
)
SELECT
gps_id,
date(gps.time),
time_grp,
dist_grp
st_setsrid(st_makeline(gps_window."position" ORDER BY gps_window.event_time), 4326) AS geom,
FROM groups
WHERE gps_time BETWEEN starttime AND endtime
GROUP BY gps.id, date(gps.time), time_grp, dist_grp
gap_id 函数正在检查来自相同 gps_id 的连续 gps 点,这些点彼此相距太远,移动速度过快或消息之间的时间过长。聚合基本上是从 gps 点创建一条线。的最终结果是一堆线,其中线中的所有点都是“合理的”。
要运行 1 天的聚合函数(开始时间 = '2020-01-01',结束时间 = '2020-01-02'),大约需要 12 秒才能完成。如果我选择一周的数据,则需要 10 分钟。如果我选择一个月的数据,则需要 15 小时以上才能完成。
我期望线性性能,因为无论如何数据都会每天分组,但事实并非如此。绕过这个性能瓶颈的明显方法是在 for 循环中运行它:
for date in date_range(starttime, endtime):
my_agg_func(date, date+1)
我可以在 Python 中做到这一点,但有什么想法可以让 for 循环在 postgres 中运行或将聚合查询更改为线性?
【问题讨论】:
-
date(gps_time)必须为每一行计算,因此 GROUP BY 操作不能利用其上的任何索引。查询太慢,无法开始。这些字段是否被索引覆盖?有多少行?在 PostgreSQL 中,您可以根据表达式创建索引,这应该会使查询速度更快 -
通常使用日历表来简化基于日期的报告。日历表每天包含一行,例如 10 到 20 年,其中包含年、月、星期、学期、季度、周数及其名称的预计算和索引字段。这样,您不必计算学期或期间的开始和结束天数,只需在日期列上与该表联接,并在所需的期间字段上进行过滤。这仍然需要在要查询的表中添加
date字段 -
TimeScaleDB 有一些用于时间序列查询的漂亮函数,但我认为在我对查询的过度优化中我停止使用它们......表大小约为每天 550 万行,并且有索引准时,gps_id,geom。
-
我将编辑查询以更符合我的实际操作。
-
gps_time上的索引没有帮助,因为查询使用了date(gps_time)的结果。尝试在date(gps_time)上创建索引
标签: sql postgresql performance postgis