【问题标题】:ST_DWITHIN not using GIST or BRIN indexST_DWITHIN 不使用 GIST 或 BRIN 索引
【发布时间】:2018-09-12 06:18:04
【问题描述】:

我正在使用 postgis 函数 ST_DWithin(geography gg1, geography gg2, double precision distance_meters) 来查找点是否在距多边形的指定距离内。我正在运行测试以查看查询需要多长时间,并且解释表明它正在对表运行顺序扫描,而不是使用 BRIN 或 GIST 索引。有人可以建议一种优化它的方法吗?

这是表格 -

table1(incident_geog) 与多边形

CREATE TABLE public.incident_geog
(
    incident_id integer NOT NULL DEFAULT nextval('incident_geog_incident_id_seq'::regclass),
    incident_name character varying(20) COLLATE pg_catalog."default",
    incident_span geography(Polygon,4326),
    CONSTRAINT incident_geog_pkey PRIMARY KEY (incident_id)
)

CREATE INDEX incident_geog_gix
    ON public.incident_geog USING gist
    (incident_span)

带有点和距离的table2(watchzones_geog)

CREATE TABLE public.watchzones_geog
(
    id integer NOT NULL DEFAULT nextval('watchzones_geog_id_seq'::regclass),
    date_created timestamp with time zone DEFAULT now(),
    latitude numeric(10,7) DEFAULT NULL::numeric,
    longitude numeric(10,7) DEFAULT NULL::numeric,
    radius integer,
    "position" geography(Point,4326),
    CONSTRAINT watchzones_geog_pkey PRIMARY KEY (id)
)

CREATE INDEX watchzones_geog_gix
    ON public.watchzones_geog USING gist
    ("position")

带有 st_dwithin 的 Sql

explain select i.incident_id,wz.id from watchzones_geog wz, incident_geog i where ST_DWithin(position,incident_span,wz.radius * 1000);

解释的输出:

Nested Loop  (cost=0.26..418436.69 rows=1 width=8)
-> Seq Scan on watchzones_geog wz  (cost=0.00..13408.01 rows=600001 width=40)
 ->  Index Scan using incident_geog_gix on incident_geog i  (cost=0.26..0.67 rows=1 width=292)
        Index Cond: (incident_span && _st_expand(wz."position", ((wz.radius * 1000))::double precision))
        Filter: ((wz."position" && _st_expand(incident_span, ((wz.radius * 1000))::double precision)) AND _st_dwithin(wz."position", incident_span, ((wz.radius * 1000))::double precision, true))

【问题讨论】:

  • 你需要输出explain analyze而不是explain并且查询有ST_DWithin(position,incident_span,50),但是查询计划似乎显示ST_DWithin(position,incident_span,(wz.radius * 1000)::double precision)
  • 存在 seq 扫描,因为您没有过滤查询中的任何内容。您要求 DB 将一张表中的所有记录(因此您进行 seq 扫描)与距另一张表 50 米的记录(因此进行索引扫描)进行比较。

标签: postgresql postgis


【解决方案1】:

您的 SQL 实际执行的是在指定距离内为 每个 点找到一些多边形。结果incident_geog.incident_id和watchzones_geog.id一一对应。因为你对每个点进行操作,所以它使用顺序扫描。

我猜你想从多边形开始找点。所以你的SQL需要改表。

explain select i.incident_id,wz.id from incident_geog i, watchzones_geog wz where ST_DWithin(position,incident_span,50);

我们可以看到:

Nested Loop  (cost=0.27..876.00 rows=1 width=16)
   ->  Seq Scan on incident_geog i  (cost=0.00..22.00 rows=1200 width=40)
   ->  Index Scan using watchzones_geog_gix on watchzones_geog wz  (cost=0.27..0.70 rows=1 width=40)
         Index Cond: ("position" && _st_expand(i.incident_span, '50'::double precision))
         Filter: ((i.incident_span && _st_expand("position", '50'::double precision)) AND _st_dwithin("position", i.incident_span, '50'::double precision, true))

因为你操作每一个订单,所以总会有一张表通过顺序扫描遍历所有记录。这两个 SQL 的结果并没有什么不同。关键是你从哪个表开始寻找另一个表的顺序。

也许你可以试试Parallel Query。不要使用Parallel Query:

SET parallel_tuple_cost TO 0;
explain analyze select i.incident_id,wz.id from incident_geog i, watchzones_geog wz where ST_DWithin(position,incident_span,50);

Nested Loop  (cost=0.27..876.00 rows=1 width=16) (actual time=0.002..0.002 rows=0 loops=1)
   ->  Seq Scan on incident_geog i  (cost=0.00..22.00 rows=1200 width=40) (actual time=0.002..0.002 rows=0 loops=1)
   ->  Index Scan using watchzones_geog_gix on watchzones_geog wz  (cost=0.27..0.70 rows=1 width=40) (never executed)
         Index Cond: ("position" && _st_expand(i.incident_span, '50'::double precision))
         Filter: ((i.incident_span && _st_expand("position", '50'::double precision)) AND _st_dwithin("position", i.incident_span, '50'::double precision, true))
 Planning time: 0.125 ms
 Execution time: 0.028 ms

尝试Parallel Query 并将parallel_tuple_cost 设置为2:

SET parallel_tuple_cost TO 2;
explain analyze select i.incident_id,wz.id from incident_geog i, watchzones_geog wz where ST_DWithin(position,incident_span,50);

Nested Loop  (cost=0.27..876.00 rows=1 width=16) (actual time=0.002..0.002 rows=0 loops=1)
       ->  Seq Scan on incident_geog i  (cost=0.00..22.00 rows=1200 width=40) (actual time=0.001..0.001 rows=0 loops=1)
       ->  Index Scan using watchzones_geog_gix on watchzones_geog wz  (cost=0.27..0.70 rows=1 width=40) (never executed)
             Index Cond: ("position" && _st_expand(i.incident_span, '50'::double precision))
             Filter: ((i.incident_span && _st_expand("position", '50'::double precision)) AND _st_dwithin("position", i.incident_span, '50'::double precision, true))
     Planning time: 0.103 ms
     Execution time: 0.013 ms

【讨论】:

  • 感谢您告诉我为什么它是对incident_geog 的seq 扫描。但是,先更改表 event_geog 再更改 watchzone_geog 的顺序并没有改善任何问题。扫描仍然需要很长时间。
  • 扫描仍然需要很长时间,因为您使用了所有记录。您可以通过设置并行工作人员的数量来尝试并行查询。
【解决方案2】:

几个一般要点:

  1. 使用 IDENTITY COLUMNS 而不是手动设置序列。
  2. 您不需要DEFAULT null::,可空列上的默认值始终为null。
  3. 确保在加载这两个表后VACUUM ANALAYZE。
  4. 不要使用 SQL-89,而是写出你的 INNER JOIN ... ON

    SELECT i.incident_id,wz.id
    FROM watchzones_geog wz
    INNER JOIN incident_geog i
      ON ST_DWithin(wz.position,i.incident_span,50);
    
  5. 在您的explain analyze 中,您的查询中有一个wz.radius * 1000,您将半径设为50。它是哪个?如果静态输入半径,查询序列是否会扫描?

  6. 如果您没有在表格中使用纬度和经度,请删除这两列。没有理由将它们存储两次。
  7. 我不会使用varchar(20),而是使用text,它更快,因为没有长度检查,而且实现方式相同。

【讨论】:

  • 嘿,埃文,感谢您的回复。第1、2、3、4、6、7点在这里没有直接关系,以提高查询执行时间。我在这里也指的是 VACUUM ANALYZE,因为我只加载表一次,而不是一直更新它们。
  • 另外,ST_DWithin 中的第三个参数是 wz.radius*1000,不适合在 EXPLAIN 和实际 sql 中发布不同的内容。
猜你喜欢
  • 2018-10-22
  • 2017-06-25
  • 1970-01-01
  • 2017-04-16
  • 1970-01-01
  • 1970-01-01
  • 2022-01-05
  • 1970-01-01
  • 2020-08-14
相关资源
最近更新 更多