【发布时间】:2013-09-25 17:15:54
【问题描述】:
我有 2 个表可以简化为这种结构:
表 1:
+----+----------+---------------------+-------+
| id | descr_id | date | value |
+----+----------+---------------------+-------+
| 1 | 1 | 2013-09-20 16:39:06 | 1 |
+----+----------+---------------------+-------+
| 2 | 2 | 2013-09-20 16:44:06 | 1 |
+----+----------+---------------------+-------+
| 3 | 3 | 2013-09-20 16:49:06 | 5 |
+----+----------+---------------------+-------+
| 4 | 4 | 2013-09-20 16:44:06 | 894 |
+----+----------+---------------------+-------+
表 2:
+----------+-------------+
| descr_id | description |
+----------+-------------+
| 1 | abc |
+----------+-------------+
| 2 | abc |
+----------+-------------+
| 3 | abc |
+----------+-------------+
| 4 | DEF |
+----------+-------------+
我想将描述加入table1,按描述过滤,所以我只得到描述= abc的行,并过滤掉“重复”行,如果两行具有相同的值并且它们的日期在6以内,则它们是重复的彼此的分钟。我想要的输出表如下,(假设 abc 是想要的描述过滤器)。
+----+----------+---------------------+-------+-------------+
| id | descr_id | date | value | description |
+----+----------+---------------------+-------+-------------+
| 1 | 1 | 2013-09-20 16:39:06 | 1 | abc |
+----+----------+---------------------+-------+-------------+
| 3 | 3 | 2013-09-20 16:49:06 | 5 | abc |
+----+----------+---------------------+-------+-------------+
我想出的查询是:
select *
from (
select *
from table1
join table2 using(descr_id)
where label='abc'
) t1
left join (
select *
from table1
join table2 using(descr_id)
where label='abc'
) t2 on( t1.date<t2.date and t1.date + interval 6 minute > t2.date)
where t1.value=t2.value.
不幸的是,这个查询需要一分钟才能运行我的数据集,并且没有返回任何结果(尽管我相信应该有结果)。有没有更有效的方法来执行这个查询?有没有办法命名派生表并稍后在同一查询中引用它?另外,为什么我的查询没有返回结果?
提前感谢您的帮助!
编辑: 我想保留几个时间戳相近的样本中的第一个。
我的 table1 有 610 万行,我的 table2 有 30K,这让我意识到 table2 只有一行用于描述“abc”。这意味着我可以事先查询 descr_id,然后使用该 id 来避免在大查询中加入 table2,从而提高效率。但是,如果我的 table2 的设置如上所述(我承认这将是糟糕的数据库设计),那么执行此类查询的好方法是什么?
【问题讨论】:
-
您是希望保留几个时间戳接近的样本中的第一个,还是最后一个,或者平均它们的时间戳,还是什么?结果集中应该使用什么时间戳来表示您的每一组样本都靠得很近?
-
好问题顺便说一句 +1 表有多少条记录?
标签: mysql performance join time-series deduplication