要从每个组返回特定行,您需要添加确定性ORDER BY 子句。基本上:
SELECT DISTINCT ON (q.id)
q.id AS question_id
, q.title AS question_title
, q.created_at AS question_created_at
, q.updated_at AS question_updated_at
, a.id AS answer_id
, a.content AS answer_content
, a.created_at AS answer_created_at
, a.updated_at AS answer_updated_at
, SUM(v.value) AS votes
FROM questions q
LEFT JOIN answers a ON q.id = a.question_id
LEFT JOIN votes v ON a.id = v.answer_id
GROUP BY q.id, a.id -- for the sum
ORDER BY q.id, a.created_at DESC NULLS LAST, a.id;
第一个ORDER BY 项必须与DISTINCT ON 子句一致。
您想要“最新”的答案,所以接下来是 a.created_at DESC。
NULLS LAST 因为该列可能为空(您没有透露)。
最后的a.id 仅在多个答案与a.created_at 并列的情况下用作决胜局。
详细解释:
加入votes 后,就不需要投票总和的相关子查询了:
(SELECT SUM(votes.value) AS votes FROM votes WHERE answers.id =votes.answer_id)
目前,您可能得到不正确的(相乘)总和。假设answers 和votes 之间存在一对多关系(否则,投票计数可以作为另一列添加到answers),它是非此即彼:要么加入表,然后GROUP BY,或者不加入表并添加相关的子查询。
我用一个简单的sum() 来修复它,假设q.id 和a.id 是它们表的各自主键(您没有透露表定义)。这是可能的,因为DISTINCT ON 在GROUP BY 之后应用。见:
或查看以下可能更好的解决方案。
当您返回所有或大多数问题时,如果您在获得最新信息后加入,则查询通常会更快每个问题的答案。喜欢:
SELECT q.id AS question_id
, q.title AS question_title
, q.created_at AS question_created_at
, q.updated_at AS question_updated_at
, a.id AS answer_id
, a.content AS answer_content
, a.created_at AS answer_created_at
, a.updated_at AS answer_updated_at
, u.user_name -- whatever you need from users table
, (SELECT SUM(value) FROM votes v WHERE v.answer_id = a.answer_id) AS votes
FROM questions q
LEFT JOIN (
SELECT DISTINCT ON (a.question_id)
a.question_id AS id
, a.id AS answer_id
, a.content AS answer_content
, a.created_at AS answer_created_at
, a.updated_at AS answer_updated_at
, a.user_id
FROM answers a
ORDER BY a.question_id, a.created_at DESC NULLS LAST, a.id
) a USING (id)
LEFT JOIN users u ON u.id = a.user_id
在这里,我保留了 votes 的相关子查询,因为这样做通常会更便宜减少选择的答案而不是计算所有答案。
users 类似(在您的答案中添加):加入 after 减少到所选答案。并将来自users 的内容放入SELECT 列表中以实际返回。
如果您的表 answers 很大,则 answer(question_id, created_at DESC NULLS LAST) 上的多列索引将是性能的理想选择。
如果每个问题有很多答案,则不同的查询技术可能会更快。见:
对于检索所有问题的一小部分,LATERAL 或相关子查询通常更快。
详细信息取决于未公开的表定义和基数。