【问题标题】:How to select value from array that is closest to value in array using vectorization?如何使用向量化从数组中选择最接近数组中的值的值?
【发布时间】:2017-04-04 09:45:57
【问题描述】:

我有一个值数组,我想从一个选项数组中替换为哪个选项是线性最接近的。

关键是选项的大小是在运行时定义的。

import numpy as np
a = np.array([[0, 0, 0], [4, 4, 4], [9, 9, 9]])
choices = np.array([1, 5, 10])

如果选择的大小是静态的,我会简单地使用 np.where

d = np.where(np.abs(a - choices[0]) > np.abs(a - choices[1]), 
      np.where(np.abs(a - choices[0]) > np.abs(a - choices[2]), choices[0], choices[2]),
         np.where(np.abs(a - choices[1]) > np.abs(a - choices[2]), choices[1], choices[2]))

获取输出:

>>d
>>[[1, 1, 1], [5, 5, 5], [10, 10, 10]]

有没有办法在保持矢量化的同时更动态地执行此操作。

【问题讨论】:

  • 从a中减去choices,求结果的最小值,代入。

标签: python performance numpy vectorization


【解决方案1】:

从a中减去选项,找到结果最小值的索引,代入。

a = np.array([[0, 0, 0], [4, 4, 4], [9, 9, 9]])
choices = np.array([1, 5, 10])
b = a[:,:,None] - choices
np.absolute(b,b)
i = np.argmin(b, axis = -1)
a = choices[i]
print a

>>> 
[[ 1  1  1]
 [ 5  5  5]
 [10 10 10]]

a = np.array([[0, 3, 0], [4, 8, 4], [9, 1, 9]])
choices = np.array([1, 5, 10])
b = a[:,:,None] - choices
np.absolute(b,b)
i = np.argmin(b, axis = -1)
a = choices[i]
print a

>>>    
[[ 1  1  1]
 [ 5 10  5]
 [10  1 10]]
>>> 

额外的维度被添加到a,以便从a 的每个元素中减去choices 的每个元素。 choices 在第三维 This link has a decent graphic 中是 broadcast 与 a。 b.shape 是 (3,3,3)。 EricsBroadcastingDoc 是一个很好的解释,最后有一个图形 3-d 示例。

第二个例子:

>>> print b
[[[ 1  5 10]
  [ 2  2  7]
  [ 1  5 10]]

 [[ 3  1  6]
  [ 7  3  2]
  [ 3  1  6]]

 [[ 8  4  1]
  [ 0  4  9]
  [ 8  4  1]]]
>>> print i
[[0 0 0]
 [1 2 1]
 [2 0 2]]
>>> 

最终分配使用Index Array 或Integer Array Indexing。

在第二个示例中,请注意元素 a[0,1] 有一个 tie,可以替换为一个或五个。

【讨论】:

  • 所以看看我是否理解这个想法是选择与 A 相同的形状?那么如果 A 是一个 10x10 的东西呢?
  • @dranobob,请参阅编辑 - 希望可以清除它。如果 a 最初是 (10,10),那么 b 本来就是 ​​(1010,3)。我们为a 创建了第三个维度,choices 可以对其进行广播。
【解决方案2】:

更详细地解释二战的excellent answer:

这个想法是创建一个新维度,它使用numpy broadcasting 将a 的每个元素与choices 中的每个元素进行比较。对于a 中的任意数量的维度,使用ellipsis syntax 很容易做到这一点:

>>> b = np.abs(a[..., np.newaxis] - choices)
array([[[ 1,  5, 10],
        [ 1,  5, 10],
        [ 1,  5, 10]],
       [[ 3,  1,  6],
        [ 3,  1,  6],
        [ 3,  1,  6]],
       [[ 8,  4,  1],
        [ 8,  4,  1],
        [ 8,  4,  1]]])

将argmin 沿您刚刚创建的轴(最后一个轴,标签为-1)为您提供choices 中您想要替换的所需索引:

>>> np.argmin(b, axis=-1)
array([[0, 0, 0],
       [1, 1, 1],
       [2, 2, 2]])

这最终允许您从choices 中选择那些元素:

>>> d = choices[np.argmin(b, axis=-1)]
>>> d
array([[ 1,  1,  1],
       [ 5,  5,  5],
       [10, 10, 10]])

对于非对称形状:

假设a 具有形状(2, 5):

>>> a = np.arange(10).reshape((2, 5))
>>> a
array([[0, 1, 2, 3, 4],
       [5, 6, 7, 8, 9]])

然后你会得到:

>>> b = np.abs(a[..., np.newaxis] - choices)
>>> b
array([[[ 1,  5, 10],
        [ 0,  4,  9],
        [ 1,  3,  8],
        [ 2,  2,  7],
        [ 3,  1,  6]],

       [[ 4,  0,  5],
        [ 5,  1,  4],
        [ 6,  2,  3],
        [ 7,  3,  2],
        [ 8,  4,  1]]])

这很难读,但它的意思是,b 有形状:

>>> b.shape
(2, 5, 3)

前两个维度来自a的形状,也就是(2, 5)。最后一个维度是您刚刚创建的维度。为了获得更好的想法:

>>> b[:, :, 0]  # = abs(a - 1)
array([[1, 0, 1, 2, 3],
       [4, 5, 6, 7, 8]])
>>> b[:, :, 1]  # = abs(a - 5)
array([[5, 4, 3, 2, 1],
       [0, 1, 2, 3, 4]])
>>> b[:, :, 2]  # = abs(a - 10)
array([[10,  9,  8,  7,  6],
       [ 5,  4,  3,  2,  1]])

注意b[:, :, i] 是a 和choices[i] 之间的绝对差,对于每个i = 1, 2, 3。

希望这有助于更清楚地解释这一点。

【讨论】:

  • 如果尺寸相差甚远,比如 A 是 17 x 7,这将如何变化?
  • @dranobob 它根本不会改变。对于非常大的数组大小,您的程序可能会占用大量内存。但是对于 17x7,您无需担心
  • @dranobob 事实上,如果您的 a 矩阵具有 (4,) 或 (12, 7, 8) 或类似的形状,它甚至不会改变,因为省略号很小心,不管a有多少个维度。
  • 这很酷,但我仍在尝试围绕矢量化概念展开思考。你能帮我想象一下这样一个奇怪的 A 矩阵在这个设置中是如何工作的吗?我可以轻松地想象 3x3
  • 我会一直盯着它,直到它变得有意义。我整个周末都在编码,所以我可能需要在早上尝试一下:)
【解决方案3】:

我喜欢broadcasting,我自己也会这样做。但是,对于大型数组,我想建议使用 np.searchsorted 的另一种方法,以保持内存效率,从而实现性能优势,就像这样 -

def searchsorted_app(a, choices):
    lidx = np.searchsorted(choices, a, 'left').clip(max=choices.size-1)
    ridx = (np.searchsorted(choices, a, 'right')-1).clip(min=0)
    cl = np.take(choices,lidx) # Or choices[lidx]
    cr = np.take(choices,ridx) # Or choices[ridx]
    mask = np.abs(a - cl) > np.abs(a - cr)
    cl[mask] = cr[mask]
    return cl

请注意,如果choices中的元素没有排序,我们需要在sorter和np.searchsorted中添加额外的参数。

运行时测试-

In [160]: # Setup inputs
     ...: a = np.random.rand(100,100)
     ...: choices = np.sort(np.random.rand(100))
     ...: 

In [161]: def broadcasting_app(a, choices): # @wwii's solution
     ...:     return choices[np.argmin(np.abs(a[:,:,None] - choices),-1)]
     ...: 

In [162]: np.allclose(broadcasting_app(a,choices),searchsorted_app(a,choices))
Out[162]: True

In [163]: %timeit broadcasting_app(a, choices)
100 loops, best of 3: 9.3 ms per loop

In [164]: %timeit searchsorted_app(a, choices)
1000 loops, best of 3: 1.78 ms per loop

相关帖子:Find elements of array one nearest to elements of array two

【讨论】:

  • 在平局(整数数组)的情况下,你喜欢 right 索引,而我喜欢 left。当我指出平局行为时,OP 没有发表评论。我猜如果数组是浮动的,那么平局的可能性就会很小。
  • 有趣,似乎性能差异随choices 的大小而变化,但与a 的大小无关。
  • @wwii 与指数有关吗?我想在choices[i] 的最后索引步骤中,选择的值将是相同的,对吧?嗯,我采用了代表 OP 在样本中的形状,即 choices.size = n 和 a.shape = (n,n)。
  • 在我的第二个示例中,a[0,1] 是三,它与两个选项(一和五)等距。这就是我所说的平局,根据 OP 的规范,可以替换一个或五个 - 两者都是正确的。你的方法选择五个,我的方法选择一个。但这可能仅在输入为整数数组时发生。我喜欢你的工作方式,我当然没有那样想的习惯。
  • @wwii 啊我现在明白了!那么 OP 需要接听这些案件,因为没有指定。
猜你喜欢
  • 2011-03-09
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-01-13
  • 1970-01-01
  • 2022-01-05
  • 1970-01-01
相关资源
最近更新 更多