【问题标题】:Is there a way to speed up this multi-table query with indices?有没有办法用索引加速这个多表查询?
【发布时间】:2019-10-04 05:27:28
【问题描述】:

我正在尝试获取属于用户所有对话的所有标签(用户通过ConversationUserPair 加入进行了许多对话) - 但查询平均需要 2,000 毫秒。

SELECT "tags"."tag_text_downcased"
FROM "tags"
INNER JOIN "conversations" ON "tags"."conversation_id" = "conversations"."id"
INNER JOIN "conversation_user_pairs" ON "conversations"."id" = "conversation_user_pairs"."conversation_id"
WHERE "conversation_user_pairs"."user_id" = ?
AND "conversation_user_pairs"."conversation_status" = ?
AND ("tags"."user_id" = ?);

当我在 psql 控制台中运行 EXPLAIN ANALYZE 时,得到的响应是:

EXPLAIN ANALYZE
SELECT "tags"."tag_text_downcased" FROM "tags" INNER JOIN "conversations" ON "tags"."conversation_id" = "conversations"."id" INNER JOIN "conversation_user_pairs" ON "conversations"."id" = "conversation_user_pairs"."conversation_id" WHERE "conversation_user_pairs"."user_id" = '459' AND "conversation_user_pairs"."conversation_status" = 'active' AND ("tags"."user_id" = '459');

Nested Loop  (cost=462.87..486.65 rows=1 width=11) (actual time=0.457..1.886 rows=40 loops=1)
   Join Filter: (tags.conversation_id = conversations.id)
   ->  Merge Join  (cost=462.78..482.97 rows=1 width=19) (actual time=0.401..1.334 rows=40 loops=1)
         Merge Cond: (tags.conversation_id = conversation_user_pairs.conversation_id)
         ->  Sort  (cost=462.70..462.83 rows=259 width=15) (actual time=0.332..0.337 rows=40 loops=1)
               Sort Key: tags.conversation_id
               Sort Method: quicksort  Memory: 27kB
               ->  Bitmap Heap Scan on tags  (cost=4.49..460.62 rows=259 width=15) (actual time=0.152..0.295 rows=40 loops=1)
                     Recheck Cond: (user_id = 459)
                     Heap Blocks: exact=23
                     ->  Bitmap Index Scan on index_tags_on_user_id_and_conversation_id  (cost=0.00..4.47 rows=259 width=0) (actual time=0.105..0.105 rows=40 loops=1)
                           Index Cond: (user_id = 459)
         ->  Index Only Scan using by_user_and_conversation_and_status on conversation_user_pairs  (cost=0.08..20.02 rows=522 width=4) (actual time=0.066..0.956 rows=390 loops=1)
               Index Cond: ((user_id = 459) AND (conversation_status = 'active'::text))
               Heap Fetches: 134
   ->  Index Only Scan using index_conversations_on_id on conversations  (cost=0.08..3.68 rows=1 width=4) (actual time=0.013..0.013 rows=1 loops=40)
         Index Cond: (id = conversation_user_pairs.conversation_id)
         Heap Fetches: 40

我认为我在有问题的三个单独的表上有适当的索引。我有:

add_index "tags", ["conversation_id", "user_id", "tag_text_downcased"], name: "find_tag_text_downcased_tags"
add_index "tags", ["conversation_id", "user_id"], name: "index_conversation_first_tags"
add_index "tags", ["user_id", "conversation_id"], name: "index_tags_on_user_id_and_conversation_id"

add_index "conversation_user_pairs", ["user_id", "conversation_id", "conversation_status"], name: "by_user_and_conversation_and_status"

add_index "conversations", ["id"], name: "index_conversations_on_id"

这里没有什么可做的来加快查询速度,因为它看起来像是在使用每个表的索引?或者有没有办法拥有一个多表索引?

【问题讨论】:

  • 您的EXPLAIN 以毫秒为单位显示时间,而不是秒。该查询的执行时间不到 2 毫秒。最慢的部分是在conversation_user_pairs 上仅扫描索引,大约需要 1 毫秒,可能是因为 134 次堆提取(表数据):explain.depesz.com/s/7DYn
  • @Ancoron 感谢您的链接,非常有帮助。是的,这个特定的查询运行得很快,但我的服务器工具显示这个查询有时会花费超过 4,000 毫秒。

标签: sql postgresql activerecord indexing postgresql-performance


【解决方案1】:

我正在用有根据的猜测来填补信息缺失...

查询

您显示的查询不适合您声明的目标:

我正在尝试获取属于用户所有对话的所有标签

“所有”是指我认为的“任何”。

还假设参照完整性是通过外键约束强制执行的。然后我们可以去掉中间人conversations。加入它只会增加成本。

按照您的方式,查询可以多次返回相同的标签。假设您想要唯一的标签,断言conversation_user_pairs 中的任何匹配行存在 就足够了。 EXISTS 半连接通常是最好的方法:

SELECT t.tag_text_downcased
FROM   tags t
WHERE  t.user_id = 459  -- assuming it's a numeric data type
AND    EXISTS (
   SELECT
   FROM   conversation_user_pairs cu
   WHERE  cu.user_id         = t.user_id
   AND    cu.conversation_id = t.conversation_id
   AND    cu.conversation_status = 'active'
   );

索引

您在tags 上的索引find_tag_text_downcased_tags 非常完美。
而by_user_and_conversation_and_status 也有好处。如果许多行不是“活动的”,而您最感兴趣的是活动的行,部分索引可能会更好:

CREATE INDEX ON conversation_user_pairs (user_id, conversation_id)
WHERE conversation_status = 'active';

这里不需要其他索引。既然你有这两个:

add_index "tags", ["conversation_id", "user_id", "tag_text_downcased"], name: "find_tag_text_downcased_tags"
add_index "tags", ["user_id", "conversation_id"], name: "index_tags_on_user_id_and_conversation_id"

...保留这个通常也没有用:

add_index "tags", ["conversation_id", "user_id"], name: "index_conversation_first_tags"

你可能会放弃它。见:

另外:如果conversation_status 只有“活动”和“死”或类似的,则将其设为boolean 列。比text更小更便宜。

【讨论】:

  • 您认为conversations 在这里没有必要是完全正确的,因为我们已经拥有conversation_id 到conversation_user_pairs。确实,我们想要“任何”(不是“全部”)标签,但tags.tag_text_downcased 实际上是一个独特的列,因此如果我理解其中的含义,“任何”和“全部”本质上是可以互换的。我从未创建过部分索引!我会阅读并尝试一下。谢谢。
  • tags.tag_text_downcased 可能是“唯一的”(究竟如何?缺少表定义),但加入conversation_user_pairs 通常会增加行数。 (看起来像 1:n 的关系。)尽管如此,EXISTS 在任何情况下都是正确的——而且通常是最快的。
猜你喜欢
  • 2014-05-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2022-11-03
  • 1970-01-01
相关资源
最近更新 更多