【问题标题】:Performance of sorted() and heapq functions in Python3Python3 中 sorted() 和 heapq 函数的性能
【发布时间】:2021-05-12 08:59:53
【问题描述】:

我想使用 Python3 以最快的方式实现以下过程:给定一个 N 随机整数列表,我需要返回 K 最小的整数(并且我不需要对返回的整数进行排序)。 我以三种不同的方式实现了它(如下面的代码所示)。

  • test_sorted() 函数使用内置的sorted() 函数对整个整数列表进行排序,然后对第一个K 元素进行切片。这个操作的代价本质上应该是运行sorted()函数的代价,它的时间复杂度是O(N log(N))。

  • test_heap() 函数使用堆来仅存储最低的K 元素并返回它们。在堆上插入一个元素的时间复杂度为O(log(N)),理论上我们需要在堆中推送一个项目的时间是N。但是,在第一次 K 插入之后,我们将从堆中推送和弹出,我希望如果传入元素大于堆中的任何元素,则不会发生插入,时间复杂度应该介于 O(K log(N)) 和O(N log(N))(取决于输入列表的实际排序)。无论如何,即使我的假设不正确,最糟糕的复杂性应该是O(N log(N))(像往常一样,我认为我们需要的所有比较的成本可以忽略不计)。

  • test_nsmallest() 函数使用来自heapq 模块的nsmallest() 函数。我对这种方法没有任何期望,因为在官方 python 文档中我只发现了

    对于较大的值,使用 sorted() 函数更有效。 我决定试一试。

# test.py

from heapq import heappush, heappushpop, nsmallest
from random import randint
from timeit import timeit

N, K = 1000, 50
RANDOM_INTS = [randint(1,100) for _ in range(N)]

def test_sorted():
    return sorted(RANDOM_INTS)[:K]

def test_heap():
    heap = []
    for val in RANDOM_INTS:
        if len(heap) < K:
            heappush(heap, -val)
        else:
            heappushpop(heap, -val)
    return [-val for val in heap]

def test_nsmallest():
    return nsmallest(K, RANDOM_INTS)


def main():
    sorted_result = timeit("test_sorted()", globals=globals(), number=100_000)
    print(f"test_sorted took: {sorted_result}")

    heap_result = timeit("test_heap()", globals=globals(), number=100_000)
    print(f"test_heap took: {heap_result}")

    nsmallest_result = timeit("test_nsmallest()", globals=globals(), number=100_000)
    print(f"test_nsmallest took: {nsmallest_result}")

    r1, r2, r3 = test_sorted(), test_heap(), test_nsmallest()
    assert len(r1) == len(r2) == len(r3)
    assert set(r1) == set(r2) == set(r3)


if __name__ == '__main__':
    main()

在我的(旧)2011 年末 MacBook Pro 上使用 2.4GHz i7 处理器的输出如下。

$ python --version
Python 3.9.2

$ python test.py 
test_sorted took: 8.389572635999999
test_heap took: 18.586762750000002
test_nsmallest took: 13.772040639000004

使用sorted() 的最简单解决方案是迄今为止最好的,谁能详细说明为什么结果不符合我的预期(即test_heap() 函数应该至少快一点)?我错过了什么?

如果我用 pypy 运行相同的代码,结果是相反的。

$ pypy --version
Python 3.7.10 (51efa818fd9b, Apr 04 2021, 12:03:51)
[PyPy 7.3.4 with GCC Apple LLVM 12.0.0 (clang-1200.0.32.29)]

$ pypy test.py 
test_sorted took: 7.1336525249998886
test_heap took: 3.1177806880004937
test_nsmallest took: 7.5453417899998385

这更接近我的期望。

假设我对 python 内部一无所知,并且我对为什么 pypy 比 python 快只有一个非常粗略的了解,任何人都可以详细说明这些结果并添加一些关于正在发生的事情的信息,以便让我正确预见未来类似情况的最佳选择?

另外,如果您对其他比上述运行速度更快的实现有任何建议,请随时分享!

更新:

如果我们需要根据某些不是项目本身值的标准对输入列表进行排序(正如我在实际用例中实际需要的那样;以上只是一个简化)?好吧,在这种情况下,结果更令人惊讶:

# test2.py

from heapq import heappush, heappushpop, nsmallest
from random import randint
from timeit import timeit


N, K = 1000, 50
RANDOM_INTS = [randint(1,100) for _ in range(N)]


def test_sorted():
    return sorted(RANDOM_INTS, key=lambda x: x)[:K]

def test_heap():
    heap = []
    for val in RANDOM_INTS:
        if len(heap) < K:
            heappush(heap, (-val, val))
        else:
            heappushpop(heap, (-val, val))
    return [val for _, val in heap]

def test_nsmallest():
    return nsmallest(K, RANDOM_INTS, key=lambda x: x)


def main():
    sorted_result = timeit("test_sorted()", globals=globals(), number=100_000)
    print(f"test_sorted took: {sorted_result}")

    heap_result = timeit("test_heap()", globals=globals(), number=100_000)
    print(f"test_heap took: {heap_result}")

    nsmallest_result = timeit("test_nsmallest()", globals=globals(), number=100_000)
    print(f"test_nsmallest took: {nsmallest_result}")

    r1, r2, r3 = test_sorted(), test_heap(), test_nsmallest()
    assert len(r1) == len(r2) == len(r3)
    assert set(r1) == set(r2) == set(r3)


if __name__ == '__main__':
    main()

哪些输出:

$ python test2.py 
test_sorted took: 18.740868524
test_heap took: 27.694126547999996
test_nsmallest took: 25.414596833000004

$ pypy test2.py 
test_sorted took: 65.88409741500072
test_heap took: 3.9442632220016094
test_nsmallest took: 19.981832798999676

这至少告诉我两件事:

  • 使用外部键进行排序非常昂贵,无论是使用 key kwarg 提供 lambda 函数,还是需要构建元组 (sorting_value, actual_value) 以获得堆中所需的排序时。

  • 将 lambdas 与 pypy 一起使用似乎非常昂贵,但我不知道为什么……也许 pypy 无法优化它们,这与它执行的其他优化不兼容???

【问题讨论】:

  • 您的test_heap 有使用直接python 的开销,而test_sorted 是用C 实现的,只有将输入参数和输出结果从/转换为python 对象的开销很小。跨度>
  • 此外,内置的sort 经过高度优化,可以处理输入中的预排序序列,考虑到您构建输入的方式,这可能很常见。

标签: python python-3.x performance time-complexity pypy


【解决方案1】:

您正在使用 CPython 解释器 和 PyPy 即时编译器 对一个小数组进行排序。结果,出现了许多复杂的开销。内置调用可能比手动编写的包含 on 循环的纯 Python 代码更快。

渐近复杂度仅适用于较大的值,因为缺少常数因素:O(n log2(n) + 30 n) 算法在实践中可能比O(2 n log2(n)) 算法慢于@987654327 @而两者都是O(n log2(n))...实际因素很难知道,因为许多重要的硬件影响应该被考虑在内。

此外,对于堆排序,所有项都必须插入堆中,这样才能得到正确的结果(不添加的可以是最小值)。这可以在O(n) 时间完成。因此,要获得n 大小列表中的第一个k 值,复杂度为O(k log(n) + n)(不考虑隐藏常量)。

使用 sorted() 的最简单解决方案是迄今为止最好的,谁能详细说明为什么结果不符合我的预期(即 test_heap() 函数至少应该快一点)?

sorted 是一个非常优化的内置函数。蟒蛇uses the very fast Timsort algorithm。 Timsort 通常比简单的 Heapsort 更快。这就是为什么它比nsmallest 快的原因,尽管它很复杂。此外,您的 Heapsort 是用纯 Python 编写的。

此外,在 CPython 中,这三种实现的大部分时间是处理排序列表和创建新列表的开销(大约是我机器上的一半时间)。 PyPy 可以减轻开销,但不能完全消除它们。 请记住,Python 列表是一个复杂的动态对象,具有许多内存间接(需要在其中存储动态类型的对象)。

假设我对 python 内部一无所知,并且我对为什么 pypy 比 python 更快只有一个非常粗略的了解,任何人都可以详细说明这些结果并添加一些关于正在发生的事情的信息,以便让我正确预见未来类似情况的最佳选择?

当您可以安全地说其中的所有类型都是本机类型时,最好的解决方案是不使用 Python 列表:固定大小的整数、简单/双精度浮点数。相反,使用 Numpy!但是,请记住,Numpy/List 转换非常缓慢。

在这里,最快的解决方案是使用np.random.randint(0, 100, N) 直接创建一个随机整数的Numpy 数组,然后使用分区算法 来检索k-使用np.partition(data, k)[:k] 的最小数字。如果需要,您可以对生成的 k 大小的数组进行排序。请注意,使用堆是执行分区的一种方法,但这远不是最快的算法(例如,请参阅QuickSelect)。最后,请注意O(n) 对RadixSort 等整数有快速排序算法。

使用 lambdas 和 pypy 似乎非常昂贵,但我不知道为什么......

AFAIK,这种情况是 PyPy (due to internal guards) 的性能问题。团队意识到了这一点,并计划在未来改进此类案例的性能。一般的经验法则是尽可能避免动态代码以获得快速执行(例如,纯 Python 对象,如 list 和 dict 以及 lambdas)。

【讨论】:

    猜你喜欢
    • 2014-08-31
    • 1970-01-01
    • 1970-01-01
    • 2018-01-04
    • 2015-08-03
    • 1970-01-01
    • 2022-10-01
    • 2012-11-28
    • 2019-01-27
    相关资源
    最近更新 更多