【问题标题】:Is this an acceptable implementation of the insertion sort algorithm?这是插入排序算法的可接受实现吗?
【发布时间】:2015-04-04 13:59:25
【问题描述】:

以下插入排序算法的 Java 实现出现在 Noel Markham 的 Java Programming Interviews Exposed 第 28 页:

public static List<Integer> insertSort(final List<Integer> numbers) {
    final List<Integer> sortedList = new LinkedList<>();
    originalList: for (Integer number : numbers) {
        for (int i = 0; i < sortedList.size(); i++) {
            if (number < sortedList.get(i)) {
                sortedList.add(i, number);
                continue originalList;
            }
        }
        sortedList.add(sortedList.size(), number);
    }
    return sortedList;
}

我的一位同事审查了此代码,发现它不能作为对“请实现插入排序算法”的面试问题的回答。他觉得数组会是排序列表更合适的数据结构。但是,正如 Markham 在同一页上解释的那样:

链表在中间添加元素非常有效 列表,只需重新排列列表中节点的指针即可。 如果使用了 ArrayList,则将元素添加到中间将是 昂贵的。 ArrayList 由数组支持,因此插入 列表的前面或中间意味着所有后续元素必须是 移动 1 到数组中的新插槽。这可以很 如果您有一个包含几百万行的列表,尤其是如果 您在列表的前面插入。

这是一个可接受的实现吗?

【问题讨论】:

  • 插入排序是就地的;你为什么要创建另一个数组?
  • 在同一页面上,Markham 说:“请注意,该方法返回一个新列表,这与冒泡排序不同,冒泡排序对元素进行就地排序。这主要是实现的选择。”但是,与您所说的插入排序的维基百科描述一致,它是“就地”:en.wikipedia.org/wiki/Insertion_sort。不知道谁说的对。

标签: java algorithm sorting insertion-sort


【解决方案1】:

考虑以下插入排序的伪代码:

for i ← 1 to length(A) - 1
    j ← i
    while j > 0 and A[j-1] > A[j]
        swap A[j] and A[j-1]
        j ← j - 1
    end while
end for

来源:http://en.wikipedia.org/wiki/Insertion_sort

1) 因此,在此过程中,您持有每个元素并将其与其前一个元素进行比较,如果前一个元素大于当前元素则交换,并且这种情况会一直发生,直到条件不满足为止。

2) 该算法的工作原理是逐个元素交换元素,而不是将元素插入到它应该在的位置。注意:-每次交换都是 o(1)。

3) 因此,在这种形式中,如果您使用列表,则需要执行 2 个操作,即连接前一个元素和当前元素,反之亦然以及相邻元素。另一方面,排序后的数组只需要一步。

4)因此,在这种方法中,排序数组比列表更有意义。

现在,如果插入排序的方法是直接将当前元素插入到合适的位置,链表会更好。

注意:- 排序后的数组或排序的链表,总体流程是一样的,只是中间步骤比排序不同。

【讨论】:

    【解决方案2】:

    从理论上讲,Markham 所说的可能是一个很好的常识:在链表中插入不应该花费太多(分配一个新节点,一些引用分配),甚至在链表末尾的插入也很便宜,因为 @ 987654321@实际上是一个双链表,并保持对最后一个元素的引用。

    插入一个新节点(对于LinkedList)和移动数组的一部分(对于ArrayList)之间的争论至少有待测试,因为ArrayList#add(int i, E)使用System.arrayCopy(),这应该是真正优化的适合这种工作。

    您可以随处听到/看到“提防微基准”。好吧,我想说的是,当您想大致了解正在发生的事情时,微基准测试可以给您一些提示...

    以下是您想要比较的 2 种方法的微基准,加上 Collections.sort() 以获得一些参考时间。注意插入排序平均是O(N^2),与Collections的tim排序比较,O(Nlog(N))。

    在插入排序的建议实现中,我只是传递了排序列表实现,以便对两个测试使用相同的函数。

    public static List<Integer> insertSort(final List<Integer> numbers,
                                           final List<Integer> sortedList) {
        //final List<Integer> sortedList = new ArrayList<>();
        originalList: for (Integer number : numbers) {
            for (int i = 0; i < sortedList.size(); i++) {
                if (number < sortedList.get(i)) {
                    sortedList.add(i, number);
                    continue originalList;
                }
            }
            sortedList.add(sortedList.size(), number);
        }
        return sortedList;
    }
    

    然后下面是一个方法,它将测量对随机整数列表进行排序并打印它所花费的时间:

    public static List<Integer> bench(List<Integer> ints, String tag,
                                      Function<List<Integer>, List<Integer>> sortf) {
        long start = System.nanoTime();
        List<Integer> sortedInts = sortf.apply(ints);
        long end = System.nanoTime();
        System.out.println(String.format("type: %6s size: %7d time(ms): %5d", 
                                         tag, ints.size(), (end-start)/1000000));
        return sortedInts;
    }
    

    函数microBench() 将循环增加数组大小,并使用 3 种方法对相同的随机数组进行排序,并比较排序后的列表。

    public static void microBench(int start, int end, int step) {
        for (int m = start; m <= end; m+=step) {
            List<Integer> ints = new Random()
                 .ints(m, 0, m).boxed()
                 .collect(Collectors.toList());
    
            List<Integer> l1 = bench(ints, "coll", (List<Integer> l) -> { 
                List<Integer> list = new ArrayList<>(l);
                Collections.sort(list);
                return list;
            });
    
            List<Integer> l2 = bench(ints, "array", (List<Integer> l) ->
                insertSort(l, new ArrayList<Integer>()));
            if (!l1.equals(l2)) {
                System.out.println("Oooops array");
            }
    
            List<Integer> l3 = bench(ints, "linked", (List<Integer> l) -> 
                insertSort(l, new LinkedList<Integer>()));
            if (!l1.equals(l3)) {
                System.out.println("Oooops linked");
            }
        }
    }
    

    所有这些都是从main 调用的。从一个 500 的数组开始,然后增加大小直到 5,000(确实非常小!)。执行环境是 MBP 2,5 GHz Intel Core i7。

    public static void main(String[] args) {
        microBench(1000, 5000, 1000);
    }
    
    type:   coll size:    1000 time(ms):     1
    type:  array size:    1000 time(ms):     8
    type: linked size:    1000 time(ms):    66
    type:   coll size:    2000 time(ms):     1
    type:  array size:    2000 time(ms):     2
    type: linked size:    2000 time(ms):   507
    type:   coll size:    3000 time(ms):     2
    type:  array size:    3000 time(ms):     4
    type: linked size:    3000 time(ms):  2283
    type:   coll size:    4000 time(ms):     1
    type:  array size:    4000 time(ms):     9
    type: linked size:    4000 time(ms):  6866
    type:   coll size:    5000 time(ms):     1
    type:  array size:    5000 time(ms):    13
    type: linked size:    5000 time(ms): 14842
    

    不用画图了解插入排序withLinkedList不是赢家! 排序 5000 个整数需要 14 秒。但是插入排序 with ArrayList 并没有那么糟糕。

    我移除了LinkedList 上的长凳,并以最大数组大小 100,000 推动了一点。

        microBench(10000, 100000, 10000);
    
    type:  array size:   10000 time(ms):    70
    type:   coll size:   20000 time(ms):     6
    type:  array size:   20000 time(ms):   290
    type:   coll size:   30000 time(ms):     8
    type:  array size:   30000 time(ms):   382
    type:   coll size:   40000 time(ms):     6
    type:  array size:   40000 time(ms):   667
    type:   coll size:   50000 time(ms):     7
    type:  array size:   50000 time(ms):   984
    type:   coll size:   60000 time(ms):     8
    type:  array size:   60000 time(ms):  1521
    type:   coll size:   70000 time(ms):    10
    type:  array size:   70000 time(ms):  2172
    type:   coll size:   80000 time(ms):    12
    type:  array size:   80000 time(ms):  2729
    type:   coll size:   90000 time(ms):    13
    type:  array size:   90000 time(ms):  3587
    type:   coll size:  100000 time(ms):    15
    type:  array size:  100000 time(ms):  4528
    

    4.5 秒与 15 毫秒。这并不奇怪,插入排序与 tim 排序/合并排序 O(NlogN) 相比仍然是 O(N^2)...

    由于 Markham 编写了大约 1,000,000 个元素数组,我只是在唯一可以正常执行的实现(来自测试的 3 个)上使用了板凳,并使用 ArrayList 删除了插入排序

        microBench(100000, 1000000, 100000);
    
    type:   coll size:  100000 time(ms):    41
    type:   coll size:  200000 time(ms):    36
    type:   coll size:  300000 time(ms):    58
    type:   coll size:  400000 time(ms):    82
    type:   coll size:  500000 time(ms):   108
    type:   coll size:  600000 time(ms):   126
    type:   coll size:  700000 time(ms):   152
    type:   coll size:  800000 time(ms):   178
    type:   coll size:  900000 time(ms):   199
    type:   coll size: 1000000 time(ms):   223
    

    223 毫秒 为 1,000,000。

    结论,当可以完成时,请注意人们可以编写和测试自己的内容! - 顺便说一句,你的同事是对的。 而且,如果您必须进行排序,插入排序通常不是要走的路。

    【讨论】:

      【解决方案3】:

      由于LinkedList 不支持RandomAccess,您必须在内循环中线性迭代链表而不是get(i)。 否则,每个元素访问都需要i 步骤来查找相应的元素,并且在将实现与ArrayList 进行比较时,另一个答案中的基准会得到误导性结果。

      public static List<Integer> insertSort(final List<Integer> numbers) {
          final LinkedList<Integer> sortedList = new LinkedList<>();
          originalList: for (Integer number : numbers) {
              int i = 0;
              for (Integer compare : sortedList) {
                  if (number < compare) {
                      sortedList.add(i, number);
                      continue originalList;
                  }
                  ++i;
              }
              sortedList.addLast(number);
          }
          return sortedList;
      }
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2016-05-17
        • 2013-08-25
        • 2020-02-22
        • 1970-01-01
        相关资源
        最近更新 更多