【问题标题】:Numba @jit(nopython=True) function offers no speed improvement on heavy Numpy functionNumba @jit(nopython=True) 函数对重型 Numpy 函数没有速度改进
【发布时间】:2019-02-14 20:29:44
【问题描述】:

我目前正在运行test_matrix_speed() 以查看我的search_and_book_availability 函数有多快。使用 PyCharm 分析器,我可以看到每个 search_and_book_availability 函数调用的平均速度为 0.001 毫秒。拥有 Numba @jit(nopython=True) 装饰器对该函数的性能没有影响。这是因为没有改进的地方而且 Numpy 在这里运行得尽可能快吗? (我不关心generate_searches函数的速度)

这是我正在运行的代码

import random

import numpy as np
from numba import jit


def generate_searches(number, sim_start, sim_end):
    searches = []
    for i in range(number):
        start_slot = random.randint(sim_start, sim_end - 1)
        end_slot = random.randint(start_slot + 1, sim_end)
        searches.append((start_slot, end_slot))
    return searches


@jit(nopython=True)
def search_and_book_availability(matrix, search_start, search_end):
    search_slice = matrix[:, search_start:search_end]
    output = np.where(np.sum(search_slice, axis=1) == 0)[0]
    number_of_bookable_vecs = output.size
    if number_of_bookable_vecs > 0:
        if number_of_bookable_vecs == 1:
            id_to_book = output[0]
        else:
            id_to_book = np.random.choice(output)
        matrix[id_to_book, search_start:search_end] = 1
        return True
    else:
        return False


def test_matrix_speed():
    shape = (10, 1440)
    matrix = np.zeros(shape)
    sim_start = 0
    sim_end = 1440
    searches = generate_searches(1000000, sim_start, sim_end)
    for i in searches:
        search_start = i[0]
        search_end = i[1]
        availability = search_and_book_availability(matrix, search_start, search_end)

【问题讨论】:

  • 随机选择预订是否重要(例如,为了避免敌对的用户输入)或者确定性地返回第一个值呢?
  • @Jatentaki - 是的。随机选择很重要

标签: python performance numpy numba


【解决方案1】:

使用您的函数和以下代码来分析速度

import time

shape = (10, 1440)
matrix = np.zeros(shape)
sim_start = 0
sim_end = 1440
searches = generate_searches(1000000, sim_start, sim_end)

def reset():
    matrix[:] = 0

def test_matrix_speed():
    for i in searches:
        search_start = i[0]
        search_end = i[1]
        availability = search_and_book_availability(matrix, search_start, search_end)

def timeit(func):
    # warmup
    reset()
    func()

    reset()
    start = time.time()
    func()
    end = time.time()

    return end - start

print(timeit(test_matrix_speed))

我发现jited 版本大约为 11.5 秒,而没有jit 则为 7.5 秒。我不是 numba 方面的专家,但它的目的是优化以非矢量化方式编写的数字代码,特别是显式的 for 循环。在您的代码中没有,您只使用矢量化操作。因此,我预计jit 的性能不会超过基线解决方案,尽管我必须承认我很惊讶看到它变得更糟。如果您希望优化您的解决方案,您可以使用以下代码缩短执行时间(至少在我的 PC 上):

def search_and_book_availability_opt(matrix, search_start, search_end):
    search_slice = matrix[:, search_start:search_end]

    # we don't need to sum in order to check if all elements are 0.
    # ndarray.any() can use short-circuiting and is therefore faster.
    # Also, we don't need the selected values from np.where, only the
    # indexes, so np.nonzero is faster
    bookable, = np.nonzero(~search_slice.any(axis=1))

    # short circuit
    if bookable.size == 0:
        return False

    # we can perform random choice even if size is 1
    id_to_book = np.random.choice(bookable)
    matrix[id_to_book, search_start:search_end] = 1
    return True

并将matrix 初始化为np.zeros(shape, dtype=np.bool),而不是默认的float64。我能够获得大约 3.8 秒的执行时间,比您的 unjited 解决方案改进了约 50%,比 jited 版本改进了约 70%。希望对您有所帮助。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-01-28
    • 2017-06-21
    • 2020-08-08
    • 1970-01-01
    • 2021-09-19
    • 1970-01-01
    • 2021-02-11
    • 1970-01-01
    相关资源
    最近更新 更多