【问题标题】:Same query, different table, different execution time on postgrespostgres 上相同的查询,不同的表,不同的执行时间
【发布时间】:2014-12-11 15:29:18
【问题描述】:

我遇到了 Postgres 的性能问题。我有两个具有相同结构、相同索引的表,并且我还在两个表的 id_coordinate 索引上执行了相同的 CLUSTER。这些表的结构如下:

     Column     |   Type   |                  Modifiers                | Storage | Description
----------------+----------+-------------------------------------------+---------+-------------
 id_best_server | integer  | not null default nextval('seq'::regclass) | plain   |
 date           | date     | not null                                  | plain   |
 id_coordinate  | integer  | not null                                  | plain   |
 mnc            | smallint |                                           | plain   |
 id_cell        | integer  |                                           | plain   |
 rx_level       | real     |                                           | plain   |
 rx_quality     | real     |                                           | plain   |
 sqi            | real     |                                           | plain   |

Indexes:
    "history_best_server_until_2013_10_pkey" PRIMARY KEY, btree (id_best_server)
    "ix_history_best_server_until_2013_10_id_coordinate" btree (id_coordinate) CLUSTER
    "ix_history_best_server_until_2013_10_id_best_server" btree (id_best_server)

查询已执行:

EXPLAIN ANALYZE SELECT DISTINCT ON (x, y) x, y, rx_level, rx_quality, date, mnc, id_cell
FROM
(
    SELECT X(co.location) AS x, Y(co.location) AS y, tems.rx_level, tems.rx_quality, date, mnc, id_cell
    FROM tems.history_best_server_until_2012_10 AS tems
    JOIN gis.coordinate AS co ON tems.id_coordinate = co.id_coordinate
        AND co.location && setsrid(makeBox2d(GeomFromText('POINT(101000 461500)', 2710),
                             GeomFromText('POINT(102400 463610)', 2710)
                             ), 2710)
    WHERE mnc = 41
) AS j1
ORDER BY x, y, date DESC

两个表的行数几乎相同(大约 8M)。当我执行上面的查询时,在一张表上我得到了这些结果:

"Unique  (cost=245742.87..245805.99 rows=8416 width=118) (actual time=3420.966..3425.584 rows=10009 loops=1)"
"  ->  Sort  (cost=245742.87..245763.91 rows=8416 width=118) (actual time=3420.963..3422.236 rows=10212 loops=1)"
"        Sort Key: (x(co.location)), (y(co.location)), tems.date"
"        Sort Method: quicksort  Memory: 1182kB"
"        ->  Hash Join  (cost=61069.15..245194.20 rows=8416 width=118) (actual time=191.365..3405.590 rows=10212 loops=1)"
"              Hash Cond: (tems.id_coordinate = co.id_coordinate)"
"              ->  Seq Scan on history_best_server_until_2012_10 tems  (cost=0.00..147705.35 rows=3226085 width=22) (actual time=0.009..1749.468 rows=3230507 loops=1)"
"                    Filter: (mnc = 41)"
"              ->  Hash  (cost=60697.73..60697.73 rows=29714 width=104) (actual time=46.828..46.828 rows=31806 loops=1)"
"                    Buckets: 4096  Batches: 1  Memory Usage: 1864kB"
"                    ->  Bitmap Heap Scan on coordinate co  (cost=937.22..60697.73 rows=29714 width=104) (actual time=14.975..35.561 rows=31806 loops=1)"
"                          Recheck Cond: (location && '0103000020960A000001000000050000000000000080A8F84000000000F02A1C410000000080A8F84000000000E84B1C41000000000000F94000000000E84B1C41000000000000F94000000000F02A1C410000000080A8F84000000000F02A1C41'::geome (...)"
"                          ->  Bitmap Index Scan on ix_coordinate_location  (cost=0.00..929.79 rows=29714 width=0) (actual time=14.593..14.593 rows=31806 loops=1)"
"                                Index Cond: (location && '0103000020960A000001000000050000000000000080A8F84000000000F02A1C410000000080A8F84000000000E84B1C41000000000000F94000000000E84B1C41000000000000F94000000000F02A1C410000000080A8F84000000000F02A1C41'::g (...)"
"Total runtime: 3426.635 ms"

在另一张桌子上,它看起来像这样:

"Unique  (cost=267070.35..267138.75 rows=9120 width=118) (actual time=172.333..177.232 rows=10051 loops=1)"
"  ->  Sort  (cost=267070.35..267093.15 rows=9120 width=118) (actual time=172.330..173.708 rows=10256 loops=1)"
"        Sort Key: (x(co.location)), (y(co.location)), tems.date"
"        Sort Method: quicksort  Memory: 1186kB"
"        ->  Nested Loop  (cost=937.22..266470.49 rows=9120 width=118) (actual time=14.876..156.322 rows=10256 loops=1)"
"              ->  Bitmap Heap Scan on coordinate co  (cost=937.22..60697.73 rows=29714 width=104) (actual time=14.788..29.510 rows=31806 loops=1)"
"                    Recheck Cond: (location && '0103000020960A000001000000050000000000000080A8F84000000000F02A1C410000000080A8F84000000000E84B1C41000000000000F94000000000E84B1C41000000000000F94000000000F02A1C410000000080A8F84000000000F02A1C41'::geometry)"
"                    ->  Bitmap Index Scan on ix_coordinate_location  (cost=0.00..929.79 rows=29714 width=0) (actual time=14.409..14.409 rows=31806 loops=1)"
"                          Index Cond: (location && '0103000020960A000001000000050000000000000080A8F84000000000F02A1C410000000080A8F84000000000E84B1C41000000000000F94000000000E84B1C41000000000000F94000000000F02A1C410000000080A8F84000000000F02A1C41'::geometr (...)"
"              ->  Index Scan using ix_history_best_server_until_2013_10_id_coordinate on history_best_server_until_2013_10 tems  (cost=0.00..6.91 rows=1 width=22) (actual time=0.003..0.003 rows=0 loops=31806)"
"                    Index Cond: (id_coordinate = co.id_coordinate)"
"                    Filter: (mnc = 41)"
"Total runtime: 178.280 ms"

总运行时间不同。

如果不使用“WHERE mnc = 41”,它们都可以快速工作。我不知道是什么导致了第一种情况下的序列扫描。请注意, mnc 只能具有 3 个可能值之一。每个值的频率在较快的表上约为 41%、39%、20%,在较慢的表上约为 43%、41%、16%。

添加: 这是快速表的统计数据。

             tablename             |    attname     | n_distinct | correlation | most_common_freqs
-----------------------------------+----------------+------------+-------------+-------------------
 history_best_server_until_2013_10 | id_best_server |         -1 |           1 |
 history_best_server_until_2013_10 | date           |       1122 |   -0.206991 | many values
 history_best_server_until_2013_10 | id_coordinate  |  -0.373645 |           1 | many values
 history_best_server_until_2013_10 | mnc            |          3 |     0.30477 | {0.411783,0.386967,0.20125}
 history_best_server_until_2013_10 | id_cell        |       5811 |  -0.0759416 | many values
 history_best_server_until_2013_10 | rx_level       |      14961 |   -0.122292 | many values
 history_best_server_until_2013_10 | rx_quality     |         16 |    0.360472 | many values
 history_best_server_until_2013_10 | sqi            |       5552 |    0.212023 | many values
(8 rows)

这个是给慢的:

             tablename             |    attname     | n_distinct | correlation | most_common_freqs
-----------------------------------+----------------+------------+-------------+-------------------
 history_best_server_until_2012_10 | id_best_server |         -1 |           1 |
 history_best_server_until_2012_10 | date           |        954 |   -0.205897 | many values
 history_best_server_until_2012_10 | id_coordinate  |  -0.421911 |           1 | many values
 history_best_server_until_2012_10 | mnc            |          3 |    0.314319 | {0.4349,0.402433,0.162667}
 history_best_server_until_2012_10 | id_cell        |       5617 |  -0.0715787 | many values
 history_best_server_until_2012_10 | rx_level       |      14129 |   -0.115288 | many values
 history_best_server_until_2012_10 | rx_quality     |         22 |    0.368943 | many values
 history_best_server_until_2012_10 | sqi            |       5320 |    0.226596 | many values

gis.coordinate 的表定义

                                                  Table "gis.coordinate"
    Column     |   Type   |                               Modifiers                                | Storage | Description
---------------+----------+------------------------------------------------------------------------+---------+-------------
 id_coordinate | integer  | not null default nextval('gis.coordinate_id_coordinate_seq'::regclass) | plain   |
 location      | geometry |                                                                        | main    |

Indexes:
    "coordinate_pkey" PRIMARY KEY, btree (id_coordinate)
    "ix_pk_coordinate" UNIQUE, btree (id_coordinate) CLUSTER
    "ix_coordinate_location" gist (location)

Check constraints:
    "enforce_dims_location" CHECK (ndims(location) = 2)
    "enforce_geotype_location" CHECK (geometrytype(location) = 'POINT'::text OR location IS NULL)
    "enforce_srid_location" CHECK (srid(location) = 2710)

【问题讨论】:

  • 我希望当你添加 where 子句时,速度较慢的会进行类型转换。
  • 不,这不是原因。我在两个表上都将 mnc 类型更改为整数,但问题仍然相同。 Select 在一张桌子上运行很快,在另一张桌子上仍然很慢。
  • 统计数据是最新的吗?如果它没有进行转换,那么它只是没有使用正确的索引。不熟悉 postgresql,但大多数都让你强制使用索引,所以也许要么更新滞后表上的统计信息,要么尝试强制索引看看它们是否解决了时间问题。
  • 能否提供gis.coordinate的表定义?并且(总是)你的 Postgres 版本。
  • 我用 gis.coordinate 和 stats 表的表定义更新了问题。我的 postgres 版本是 9.1.13。

标签: sql database performance postgresql postgis


【解决方案1】:

这不是相同的数据,因此期望相同的计划是不合理的,除非统计信息(mnc = 41 的行数等,值在整个表中的分布方式等)相似。

在一种情况下,该值很可能是频繁的并遍布各处,而在另一种情况下则被狭窄地分组。在第一种情况下,对行进行 seq 扫描通常会更快;另一方面,索引扫描通常会更快。

【讨论】:

  • 我用两个表的统计表更新了问题。也许我错了,但我认为统计数据足够相似,可以使用相同的查询计划。
  • 是的,所以频率或多或少相似(假设从高到低的百分比表示相同的值)但它仍然可能与数据分布或类似的东西有关。我的预感(如果有的话)是@ErwinBrandstetter 发布的出色答案在建议调查您当前的聚集索引是否有效和有用方面是正确的——但是,请确保将其替换为对您正在运行的其他查询有用的东西在桌子上。
【解决方案2】:

从表coordinate获取行的计划在这两种情况下是相同的。

在对表 2012_10 的较慢查询中,Postgres 在 seq 扫描中直接从表中读取行,并将哈希连接到来自 coordinate 的结果。

在对表 2013_10 的更快查询中,Postgres 从id_coordinate 上的索引中收集行,按mnc 过滤结果并使用来自coordinate 的结果运行嵌套循环。

显然 Postgres 希望索引为第二次查询付费。快速情况下的 41% mnc = 41 比慢速情况下的 43% 更具选择性。通常不足以解释差异,但临界点就在某个地方。显然,切换到 seqscan 的决定是一个糟糕的决定,所以我猜你的成本设置应该进行调整。见下文。此外,我会使用 mnc 的频率较低的值进行测试。

许多细节会影响决策和结果:

  • 值的频率和分布。
  • 死元组导致的表膨胀(您使用 CLUSTER 删除了它,因此我们可以排除这种情况)。
  • 缓存大小以及索引和/或表是否已缓存(因为之前从该会话或其他会话调用了同一个表)。在测试中多次运行每个查询以平衡竞争环境。
  • Statistics Used by the Planner,由ANALYZE 收集,以对情况进行现实估计。
  • Planner Cost Constants. 如果您的统计数据不是问题,那么您应该在这里调整设置。
  • 配置设置,尤其是shared_buffers、work_mem 和effective_cache_size。

CLUSTER

Per documentation:

因为规划器记录了关于表排序的统计信息, 建议在新聚集的表上运行ANALYZE。 否则,规划器可能会做出糟糕的查询计划选择。

我的大胆强调。这可能是(部分)答案。

id_coordinate 上的 CLUSTER 也无济于事。出于查询的目的,似乎没有改善行的局部性。我建议你创建一个额外的多列索引

CREATE INDEX ix_history_?? ON history_?? (mnc, id_coordinate);

还有

CLUSTER history_?? USING ix_history_??;

这应该会有所帮助 - 对于组合索引扫描而不是 filter 步骤,索引也应该更快。

在当前版本的 Postgres 中更好的索引

没有解释这种现象,但是您的 Postgres 版本 9.1.13 已经过时并且是一个限制因素。自您的版本以来的许多改进。特别是对于大数据和索引。

使用 pg 9.2+,您可以从仅索引扫描中获益:在 gis.coordinate 的 GiST 索引中包含 id_coordinate 以使其成为多列索引:

CREATE INDEX ix_coordinate_location ON gis.coordinate (id_coordinate, location)

为此,您需要额外的模块 btree_gist。详情:

更简单的查询

无论哪种方式,您都可以简化查询:

SELECT DISTINCT ON (1, 2)
       X(co.location) AS x
     , Y(co.location) AS y
     , tems.rx_level, tems.rx_quality, date, mnc, id_cell
FROM   tems.history_best_server_until_2012_10 AS tems
JOIN   gis.coordinate AS co USING (id_coordinate)
WHERE  co.location
    && setsrid(makeBox2d(GeomFromText('POINT(101000 461500)', 2710)
                       , GeomFromText('POINT(102400 463610)', 2710)), 2710)
AND    mnc = 41
ORDER  BY 1, 2, date DESC;

虽然没有子查询,但我预计不会对性能产生太大影响。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2017-01-28
    • 2021-09-23
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-08-19
    • 1970-01-01
    相关资源
    最近更新 更多