【发布时间】:2021-12-19 13:45:16
【问题描述】:
我有以下单元格:
cells = np.array([[1, 1, 1],
[1, 1, 0],
[1, 0, 0],
[1, 0, 1],
[1, 0, 0],
[1, 1, 1]])
我想计算水平和垂直邻接来得出这个结果:
# horizontal adjacency
array([[3, 2, 1],
[2, 1, 0],
[1, 0, 0],
[1, 0, 1],
[1, 0, 0],
[3, 2, 1]])
# vertical adjacency
array([[6, 2, 1],
[5, 1, 0],
[4, 0, 0],
[3, 0, 1],
[2, 0, 0],
[1, 1, 1]])
实际的解决方案是这样的:
def get_horizontal_adjacency(cells):
adjacency_horizontal = np.zeros(cells.shape, dtype=int)
for y in range(cells.shape[0]):
span = 0
for x in reversed(range(cells.shape[1])):
if cells[y, x] > 0:
span += 1
else:
span = 0
adjacency_horizontal[y, x] = span
return adjacency_horizontal
def get_vertical_adjacency(cells):
adjacency_vertical = np.zeros(cells.shape, dtype=int)
for x in range(cells.shape[1]):
span = 0
for y in reversed(range(cells.shape[0])):
if cells[y, x] > 0:
span += 1
else:
span = 0
adjacency_vertical[y, x] = span
return adjacency_vertical
算法基本上是(对于水平邻接):
- 循环通过行
- 通过列向后循环
- 如果单元格的 x、y 值不为零,则在实际跨度上加 1
- 如果单元格的 x、y 值为为零,则将实际跨度重置为零
- 将跨度设置为结果数组的新 x、y 值
由于我需要在所有数组元素上循环两次,这对于较大的数组(例如图像)来说很慢。
有没有办法使用矢量化或其他一些 numpy 魔法来改进算法?
总结:
joni 和 Mark Setchell 提出了很好的建议!
我创建了一个带有示例图像的small Repo 和一个带有比较的python 文件。结果令人惊讶:
- 原始方法:3.675 秒
- 使用 Numba:0.002 秒
- 使用 Cython:0.005 秒
【问题讨论】:
-
不,我想在 Python 中实现 evryway.com/largest-interior
-
恕我直言,加速函数的最简单方法是在 Cython 中重写两者。它或多或少是相同的代码。
标签: python numpy performance optimization vectorization