【问题标题】:Postgres: Slow query when using OR statement in a join queryPostgres:在连接查询中使用 OR 语句时查询速度慢
【发布时间】:2021-04-10 04:31:23
【问题描述】:

我们在 2 个表之间运行连接查询。 该查询有一个 OR 语句,用于比较左表中的一列和右表中的一列。查询性能非常低,我们通过将 OR 更改为 UNION 来修复它。

为什么会这样?我正在寻找可能阐明该问题的详细解释或文档参考。


使用 Or 语句查询:

db1=# explain analyze select count(*) 
from conversations 
join agents on conversations.agent_id=agents.id 
where conversations.id=1 or agents.id = '123';

**Query plan**                                                                       
----------------------------------------------------------------------------------------------------------------------------------
 Finalize Aggregate  (cost=**11017.95..11017.96** rows=1 width=8) (actual time=54.088..54.088 rows=1 loops=1)
   ->  Gather  (cost=11017.73..11017.94 rows=2 width=8) (actual time=53.945..57.181 rows=3 loops=1)
        Workers Planned: 2
        Workers Launched: 2
        ->  Partial Aggregate  (cost=10017.73..10017.74 rows=1 width=8) (actual time=48.303..48.303 rows=1 loops=3)
            ->  Hash Join  (cost=219.26..10016.69 rows=415 width=0) (actual time=5.292..48.287 rows=130 loops=3)
                    Hash Cond: (conversations.agent_id = agents.id)
                    Join Filter: ((conversations.id = 1) OR ((agents.id)::text = '123'::text))
                    Rows Removed by Join Filter: 80035
                    ->  Parallel Seq Scan on conversations  (cost=0.00..9366.95 rows=163995 width=8) (actual time=0.017..14.972 rows=131196 loops=3)
                    ->  Hash  (cost=143.56..143.56 rows=6056 width=16) (actual time=2.686..2.686 rows=6057 loops=3)
                        Buckets: 8192  Batches: 1  Memory Usage: 353kB
                        ->  Seq Scan on agents  (cost=0.00..143.56 rows=6056 width=16) (actual time=0.011..1.305 rows=6057 loops=3)
 Planning time: 0.710 ms
 Execution time: 57.276 ms
(15 rows)

将 OR 更改为 UNION:

db1=# explain analyze select count(*) from (
  select * 
  from conversations 
    join agents on conversations.agent_id=agents.id 
  where conversations.installation_id=1 
  union 
  select * 
  from conversations 
    join agents on conversations.agent_id=agents.id 
  where agents.source_id = '123') as subquery;
                                                   
**Query plan:**
               
----------------------------------------------------------------------------------------------------------------------------------
 Aggregate  (**cost=1114.31..1114.32** rows=1 width=8) (actual time=8.038..8.038 rows=1 loops=1)
   ->  HashAggregate  (cost=1091.90..1101.86 rows=996 width=1437) (actual time=7.783..8.009 rows=390 loops=1)
        Group Key: conversations.id, conversations.created, conversations.modified, conversations.source_created, conversations.source_id, conversations.installation_id, bra
in_conversation.resolution_reason, conversations.solve_time, conversations.agent_id, conversations.submission_reason, conversations.is_marked_as_duplicate, conversations.n
um_back_and_forths, conversations.is_closed, conversations.is_solved, conversations.conversation_type, conversations.related_ticket_source_id, conversations.channel, brain_convers
ation.last_updated_from_platform, conversations.csat, agents.id, agents.created, agents.modified, agents.name, agents.source_id, organizati
on_agent.installation_id, agents.settings
        ->  Append  (cost=219.68..1027.16 rows=996 width=1437) (actual time=5.517..6.307 rows=390 loops=1)
            ->  Hash Join  (cost=219.68..649.69 rows=931 width=224) (actual time=5.516..6.063 rows=390 loops=1)
                    Hash Cond: (conversations.agent_id = agents.id)
                    ->  Index Scan using conversations_installation_id_b3ff5c00 on conversations  (cost=0.42..427.98 rows=931 width=154) (actual time=0.039..0.344 rows=879 loops=1)
                        Index Cond: (installation_id = 1)
                    ->  Hash  (cost=143.56..143.56 rows=6056 width=70) (actual time=5.394..5.394 rows=6057 loops=1)
                        Buckets: 8192  Batches: 1  Memory Usage: 710kB
                        ->  Seq Scan on agents  (cost=0.00..143.56 rows=6056 width=70) (actual time=0.014..1.938 rows=6057 loops=1)
            ->  Nested Loop  (cost=0.70..367.52 rows=65 width=224) (actual time=0.210..0.211 rows=0 loops=1)
                    ->  Index Scan using agents_source_id_106c8103_like on agents agents_1  (cost=0.28..8.30 rows=1 width=70) (actual time=0.210..0.210 rows=0 loops=1)
                        Index Cond: ((source_id)::text = '123'::text)
                    ->  Index Scan using conversations_agent_id_de76554b on conversations conversations_1  (cost=0.42..358.12 rows=110 width=154) (never executed)
                        Index Cond: (agent_id = agents_1.id)
 Planning time: 2.024 ms
 Execution time: 9.367 ms
(18 rows)

【问题讨论】:

  • 你能试试这个吗?: select count(*) from conversations,agents where conversations.agent_id=agents.id and ( conversations.id=1 or agent.id= '123') ;您需要在行中的索引 conversations.agent_id , conversations.id , agents.id
  • 你认为这些说法是一样的吗? OR 条件有 conversations.id 和 agent.id 但 UNION 有 conversations.installation_id 和 agent.source_id ?
  • @NevilleKuyt 15 和 18 是执行计划本身中的行,而不是底层查询返回的结果。这些结果也可能不同,但您无法从查询计划中判断它们是否存在。

标签: sql postgresql join query-optimization


【解决方案1】:

是的。 or 有一种杀死查询性能的方法。对于这个查询:

select count(*) 
from conversations c join
     agents a
     on c.agent_id = a.id 
where c.id = 1 or a.id = 123;

请注意,我删除了 123 周围的引号。它看起来像一个数字,所以我认为它是。对于此查询,您需要在conversations(agent_id) 上建立索引。

编写查询的最有效方法可能是:

select count(*)
from ((select 1
       from conversations c join
            agents a
            on c.agent_id = a.id 
       where c.id = 1
      ) union all
      (select 1
       from conversations c join
            agents a
            on c.agent_id = a.id 
       where a.id = 123 and c.id <> 1
      )
     ) ac;

注意使用union all 而不是union。附加的where 条件消除了重复。

这可以利用以下索引:

  • conversations(id, agent_id)
  • agents(id)
  • conversations(agent_id, id)

【讨论】:

    猜你喜欢
    • 2021-10-09
    • 2015-03-22
    • 1970-01-01
    • 2023-03-22
    • 2021-08-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多