【问题标题】:Using MySQL Index Scans for Sorts使用 MySQL 索引扫描进行排序
【发布时间】:2014-07-29 02:07:48
【问题描述】:

我通过“高性能 MySQL”一书学习 MySQL 索引详细信息,但我无法理解一件事。

正如书中所说,(第 124 页使用索引扫描进行排序)

MySQL 有两种方法来产生有序的结果:它可以使用文件排序, 或者它可以按顺序扫描索引。

仅当索引的顺序为 与 ORDER BY 子句完全相同,所有列都按 相同的方向(上升或下降)。

ORDER BY 子句也具有与查找查询相同的限制:它 需要形成索引的最左前缀。在所有其他情况下, MySQL 使用文件排序。

此外,作者给出了一些使用 MySQL Sakila 的示例数据库的示例 [http://dev.mysql.com/doc/sakila/en/][1]

第一个例子运行良好:

标准 Sakila 示例数据库中的出租表有一个索引 (rental_date、inventory_id、customer_id):

CREATE TABLE rental (
...
PRIMARY KEY (rental_id),
UNIQUE KEY rental_date (rental_date,inventory_id,customer_id),
KEY idx_fk_inventory_id (inventory_id),
KEY idx_fk_customer_id (customer_id),
KEY idx_fk_staff_id (staff_id),
...
);

MySQL 使用 rent_date 索引来排序以下查询,因为您 从 EXPLAIN 中缺少文件排序可以看出:

> mysql> EXPLAIN SELECT
> rental_id, staff_id FROM sakila.rental
> -> WHERE rental_date = '2005-05-25'
> -> ORDER BY inventory_id, customer_id\G
> *************************** 1. row *************************** 
> type: ref 
> possible_keys: rental_date 
> key: rental_date 
> rows: 1 
> Extra: Using where 

即使 ORDER BY 子句本身不是 索引的最左边前缀,因为我们指定了相等 索引中第一列的条件。

需要注意的是:它们在 where 子句中使用索引列,但在 SELECT 查询中使用不同的列。

第二个例子简明扼要:

以下查询也有效,因为 ORDER BY 中的两列 是索引的最左侧前缀:

... WHERE rent_date > '2005-05-25' ORDER BY rent_date,inventory_id;

但在这里您可以获得不同的结果,而不是您的 SELECT 列内容:

第一种情况,使用文件排序:

EXPLAIN 
SELECT `rental_id`, `staff_id` FROM `sakila`.`rental`
WHERE `rental_date` > '2005-05-25'
ORDER BY `rental_date`, `inventory_id`;

类型:全部 可能的密钥:出租日期
键:空 额外:使用where;使用文件排序

第二种情况,使用索引:

EXPLAIN 
SELECT `rental_id`, `rental_date`, `inventory_id` FROM `sakila`.`rental`
WHERE `rental_date` > '2005-05-25'
ORDER BY `rental_date`, `inventory_id`;

类型:范围 possible_key:出租日期 键:rental_date 额外:使用where;使用索引

为什么它会以这种奇怪的方式工作?如前所示,第一个示例使用索引排序,即使在 SELECT 子句中使用 WHERE 子句包含不同的列。

【问题讨论】:

    标签: mysql sorting indexing


    【解决方案1】:

    在第二个查询中:

    SELECT `rental_id`, `rental_date`, `inventory_id` FROM `sakila`.`rental`
    WHERE `rental_date` > '2005-05-25'
    ORDER BY `rental_date`, `inventory_id`;
    

    MySql 直接从索引中获取数据,根本不引用表。
    请查看索引定义并将其与查询引用的列进行比较:

    UNIQUE KEY rental_date (rental_date,inventory_id,customer_id)
    

    索引包含查询引用的所有列,但 rental_id 除外,但 rental_id 是主键,除了在其定义中明确给出的列之外,每个索引始终还包含主键值。
    这是此查询的覆盖索引,请参见此处:http://en.wikipedia.org/wiki/Index_%28database%29#Covering_index


    但是在第一个查询中:

    SELECT `rental_id`, `staff_id` FROM `sakila`.`rental`
    WHERE `rental_date` > '2005-05-25'
    ORDER BY `rental_date`, `inventory_id`;
    

    有staff_id 列,未存储在索引中。
    在这种情况下,MySql 必须首先检索匹配 WHERE 条件的索引条目,然后对于每个条目必须从表中获取整条记录以获取该条目缺少的 staff_id 值。


    现在请对您的数据库运行此查询并检查其结果:

    select count(*) As total,
               sum( case when `rental_date` > '2005-05-25' then 1 else 0 end ) As x1,
               sum( case when `rental_date` = '2005-05-25' then 1 else 0 end ) As x0
    from rental
    ;
    

    在我的sakila database 副本中,此查询返回以下内容:

    + ---------- + ------- + ------- +
    | total      | x1      | x0      |
    + ---------- + ------- + ------- +
    | 16044      | 16036   | 0       |
    + ---------- + ------- + ------- +
    

    如您所见,表中几乎所有的记录(99.9%)都大于2005-05-25。 在这种情况下,MySql 决定不使用索引从表中检索行,而是更愿意将表的全部内容加载到内存中,并在这里进行排序 - 表相对较小,仅包含 16k 条记录。
    但是,如果您恢复条件,MySql 更喜欢索引访问方法:

    EXPLAIN 
    SELECT `rental_id`, `staff_id` FROM `sakila`.`rental`
    WHERE `rental_date` < '2005-05-25'
    ORDER BY `rental_date`, `inventory_id`;
    + ------- + ---------------- + ---------- + --------- + ------------------ + -------- + ------------ + -------- + --------- + ---------- +
    | id      | select_type      | table      | type      | possible_keys      | key      | key_len      | ref      | rows      | Extra      |
    + ------- + ---------------- + ---------- + --------- + ------------------ + -------- + ------------ + -------- + --------- + ---------- +
    | 1       | SIMPLE           | rental     | range     | rental_date        | rental_date | 5            |          | 8         | Using index condition |
    + ------- + ---------------- + ---------- + --------- + ------------------ + -------- + ------------ + -------- + --------- + ---------- +
    

    为什么它在第一种情况下不使用索引?因为使用索引条目从表中检索记录通常是最昂贵的方法 - 真的 :)
    该索引仅在必须检索表的很小部分的情况下才有效 - 几个百分比,可能 要从id 的表中仅获取一条记录,MySql 必须获取包含多条记录的整页(块)数据。使用排序索引,我们必须从表中的不同位置一一获取记录,因此当我们想使用索引获取表的 90% 时,需要多次检索相同的数据块。在这种情况下,只按顺序读取它们一次并在内存中排序会更容易也更便宜。

    【讨论】:

      猜你喜欢
      • 2013-03-07
      • 2018-04-15
      • 2020-12-22
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-10-15
      • 1970-01-01
      • 2014-11-27
      相关资源
      最近更新 更多