【发布时间】:2009-09-23 20:17:33
【问题描述】:
我有办法在 Postgres 中强制执行特定的加入顺序吗?
我有一个看起来像这样的查询。我已经消除了实际查询中的一堆东西,但是这种简化说明了这个问题。剩下的不应该太神秘:使用角色/任务安全系统,我试图确定给定用户是否有权执行给定任务。
select task.taskid
from userlogin
join userrole using (userloginid)
join roletask using (roleid)
join task using (taskid)
where loginname='foobar'
and taskfunction='plugh'
但我意识到程序已经知道 userlogin 的值,所以似乎可以通过跳过对 userlogin 的查找并只填写 userloginid 来提高查询效率,如下所示:
select task.taskid
from userrole
join roletask using (roleid)
join task using (taskid)
where userloginid=42
and taskfunction='plugh'
当我这样做时——从查询中删除一个表并硬编码从该表中检索到的值——解释计划时间增加了!在原始查询中,Postgres 读取 userlogin 然后 userrole 然后 roletask 然后 task。但在新查询中,它决定先读取 roletask,然后加入 userrole,尽管这需要对 roletask 进行全文件扫描。
完整的解释计划是:
版本 1:
Hash Join (cost=12.79..140.82 rows=1 width=8)
Hash Cond: (roletask.taskid = task.taskid)
-> Nested Loop (cost=4.51..129.73 rows=748 width=8)
-> Nested Loop (cost=4.51..101.09 rows=12 width=8)
-> Index Scan using idx_userlogin_loginname on userlogin (cost=0.00..8.27 rows=1 width=8)
Index Cond: ((loginname)::text = 'foobar'::text)
-> Bitmap Heap Scan on userrole (cost=4.51..92.41 rows=33 width=16)
Recheck Cond: (userrole.userloginid = userlogin.userloginid)
-> Bitmap Index Scan on idx_userrole_login (cost=0.00..4.50 rows=33 width=0)
Index Cond: (userrole.userloginid = userlogin.userloginid)
-> Index Scan using idx_roletask_role on roletask (cost=0.00..1.50 rows=71 width=16)
Index Cond: (roletask.roleid = userrole.roleid)
-> Hash (cost=8.27..8.27 rows=1 width=8)
-> Index Scan using idx_task_taskfunction on task (cost=0.00..8.27 rows=1 width=8)
Index Cond: ((taskfunction)::text = 'plugh'::text)
版本 2:
Hash Join (cost=96.58..192.82 rows=4 width=8)
Hash Cond: (roletask.roleid = userrole.roleid)
-> Hash Join (cost=8.28..104.10 rows=9 width=16)
Hash Cond: (roletask.taskid = task.taskid)
-> Seq Scan on roletask (cost=0.00..78.35 rows=4635 width=16)
-> Hash (cost=8.27..8.27 rows=1 width=8)
-> Index Scan using idx_task_taskfunction on task (cost=0.00..8.27 rows=1 width=8)
Index Cond: ((taskfunction)::text = 'plugh'::text)
-> Hash (cost=87.92..87.92 rows=31 width=8)
-> Bitmap Heap Scan on userrole (cost=4.49..87.92 rows=31 width=8)
Recheck Cond: (userloginid = 42)
-> Bitmap Index Scan on idx_userrole_login (cost=0.00..4.49 rows=31 width=0)
Index Cond: (userloginid = 42)
(是的,我知道在这两种情况下成本都很低,而且差异看起来并不重要。但这是在我从查询中消除了一堆额外的工作以简化我必须发布的内容之后。真正的查询仍然不离谱,但我对原理更感兴趣。)
【问题讨论】:
-
你能显示查询计划(解释分析)和表定义吗?
-
好吧,你问了,我把假设的简单示例替换为真实的查询,并添加了解释计划结果。哦,我确信我可以添加一些额外的索引来加快第二次查询,但这不是重点。考虑到 Postgres 的查询,为什么选择一个不如它所能做的最好的计划?特别是当它证明如果我使查询更复杂它可以做得更好?
-
您是在比较查询的实际运行时间还是只查看解释计划中的标题成本?花费?值得注意的是,您刚刚发布了解释输出,而不是解释分析。尽管您预计更高的成本等同于较慢的查询运行,但可能不会这样。
-
没错,我没有比较实际的运行时间,你当然是正确的,相对速度可能会逆转。但我认为这无关紧要:Postgres 必须根据计算的计划成本而不是实际成本来制定计划决策,并且按照它自己的标准,它选择了一个劣质计划。如果实际运行时间变得更好,那只是运气不好。
标签: sql database postgresql