【问题标题】:Joining two tables and then fetching n% of the rows randomly from the result. (Querying the tables in http://data.stackexchange.com/ )连接两个表,然后从结果中随机获取 n% 的行。 (查询 http://data.stackexchange.com/ 中的表)
【发布时间】:2016-05-16 03:04:07
【问题描述】:
在从 stackoverflow 数据转储 (https://data.stackexchange.com/) 中加入两个表(用户和帖子)后,我试图获取 1% 的结果行的随机样本。
我使用了以下查询:
select top 1 percent * from users u join posts p ON p.OwnerUserId = u.Id
order by newid();
由于某些服务器对执行时间的限制,我收到了错误:
错误:“超时已过。在操作完成之前超时时间已过或服务器没有响应。”
有人可以建议我如何优化查询吗?
【问题讨论】:
标签:
sql-server
optimization
【解决方案1】:
从大表中选择随机数据时,newid() 并不是一个真正的好选择,因为它需要对所有行进行排序 -- 如果只选择 1%,那会浪费很多时间。
微软推荐使用binary_checksum到select rows randomly,如果1%的准确率不重要,这应该会好很多:
select * from Users u
join (
select * from Posts
WHERE (ABS(CAST(
(BINARY_CHECKSUM
(Id, NEWID())) as int))
% 100) < 1
) p on p.OwnerUserId = u.Id
由于帖子是视图,因此无法使用tablesample,但在实际情况下也可以选择。
【解决方案2】:
使用rand(),您可以通过这种方式显示随机数量的行:
set @r = rand();
SELECT * FROM `anuncios` WHERE rand() < @r
- 请注意,如果您想获得最小、最大甚至特定百分比的播放记录,请将
rvariable 设置为您需要的任何值。