【问题标题】:How can I calculate distance between points in each row of an array如何计算数组每一行中点之间的距离
【发布时间】:2022-01-17 18:45:02
【问题描述】:

我有一个这样的数组,我必须找到每个点之间的距离。我怎么能用 numpy 在 python 中这样做?

array([[  8139, 112607],
       [  8139, 115665],
       [  8132, 126563],
       [  8193, 113938],
       [  8193, 123714],
       [  8156, 120291],
       [  8373, 125253],
       [  8400, 131442],
       [  8400, 136354],
       [  8401, 129352],
       [  8439, 129909],
       [  8430, 135706],
       [  8430, 146359],
       [  8429, 139089],
       [  8429, 133243]])

【问题讨论】:

标签: python arrays pandas numpy


【解决方案1】:

让我们把这个问题减少到 4 个点:

points = np.array([[8139, 115665], [8132, 126563], [8193, 113938], [8193, 123714]])

一般来说,你需要做2个步骤:

  • 为您想要获取的点对创建索引
  • 为这些配对申请 np.hypot。

TL;DR

制作点索引

您可以通过多种方式为要获取的每对点创建索引对。但它们来自哪里?在任何情况下,从邻接矩阵开始构建它们都是一个好主意。

案例 1

以最常见的方式,您可以像这样开始构建它:

adjacency = np.ones(shape=(len(points), len(points)), dtype=bool)
>>> adjacency
[[ True  True  True  True]
 [ True  True  True  True]
 [ True  True  True  True]
 [ True  True  True  True]]

它对应于您需要采取的索引,如下所示:

adjacency_idx_view = np.transpose(np.nonzero(adjacency))
for n in adjacency_idx_view.reshape(len(points), len(points), 2):
>>> print(n.tolist())
[[0, 0], [1, 0], [2, 0], [3, 0]]
[[0, 1], [1, 1], [2, 1], [3, 1]]
[[0, 2], [1, 2], [2, 2], [3, 2]]
[[0, 3], [1, 3], [2, 3], [3, 3]]

这就是你收集它们的方式:

x, y = np.nonzero(adjacency)
>>> np.transpose([x, y])
array([[0, 0],
       [0, 1],
       [0, 2],
       [0, 3],
       [1, 0],
       [1, 1],
       [1, 2],
       [1, 3],
       [2, 0],
       [2, 1],
       [2, 2],
       [2, 3],
       [3, 0],
       [3, 1],
       [3, 2],
       [3, 3]], dtype=int64)

也可以像 @ 一样手动完成 Corralien 的回答:

x = np.repeat(np.arange(len(points)), len(points))
y = np.tile(np.arange(len(points)), len(points))

案例 2

在以前的情况下,每对点都是重复的。还有点重复的对。更好的选择是忽略这些过多的数据,只取第一个点的索引小于第二个点的索引的对:

adjacency = np.less.outer(np.arange(len(points)), np.arange(len(points)))
>>> print(adjacency)
[[False  True  True  True]
 [False False  True  True]
 [False False False  True]
 [False False False False]]
x, y = np.nonzero(adjacency)

这并没有被广泛使用。尽管这超出了np.triu_indices 的范围。因此,作为替代方案,我们可以使用:

x, y = np.triu_indices(len(points), 1)

这会导致:

>>> np.transpose([x, y])
array([[0, 1],
       [0, 2],
       [0, 3],
       [0, 4],
       [1, 2],
       [1, 3],
       [1, 4],
       [2, 3],
       [2, 4],
       [3, 4]])

案例 3 您也可以尝试仅省略成对的重复点,并留下点被交换的对。与 案例 1 一样,它需要 2 倍的内存和消耗时间,所以我将其仅用于演示目的:

adjacency = ~np.identity(len(points), dtype=bool)
>>> adjacency
array([[False,  True,  True,  True],
       [ True, False,  True,  True],
       [ True,  True, False,  True],
       [ True,  True,  True, False]])
x, y = np.nonzero(adjacency)
>>> np.transpose([x, y])
array([[0, 1],
       [0, 2],
       [0, 3],
       [1, 0],
       [1, 2],
       [1, 3],
       [2, 0],
       [2, 1],
       [2, 3],
       [3, 0],
       [3, 1],
       [3, 2]], dtype=int64)

我将手动制作x 和y(不加掩码)作为其他人的练习。

申请np.hypot

您可以使用np.hypot(np.transpose(a - b)) 代替np.sqrt(np.sum((a - b) ** 2, axis=1))。我将把我的 案例 2 作为我的索引生成器:

def distance(points):
    x, y = np.triu_indices(len(points), 1)
    x_coord, y_coord = np.transpose(points[x] - points[y])
    return np.hypot(x_coord, y_coord)

>>> distance(points)
array([10898.00224812,  1727.84403231,  8049.18113848, 12625.14736548,
        2849.65296133,  9776.        ])

【讨论】:

    【解决方案2】:

    您可以使用np.repeat 和np.tile 创建所有组合,然后计算欧几里得距离:

    xy = np.array([[8139, 115665], [8132, 126563], [8193, 113938], [8193, 123714],
                   [8156, 120291], [8373, 125253], [8400, 131442], [8400, 136354],
                   [8401, 129352], [8439, 129909], [8430, 135706], [8430, 146359],
                   [8429, 139089], [8429, 133243]])
    
    a = np.repeat(xy, len(xy), axis=0)
    b = np.tile(xy, [len(xy), 1])
    d = np.sqrt(np.sum((a - b) ** 2, axis=1))
    

    d 的输出是 (196,),即 14 x 14。

    更新

    但我必须在函数中完成。

    def distance(xy):
        a = np.repeat(xy, len(xy), axis=0)
        b = np.tile(xy, [len(xy), 1])
        return np.sqrt(np.sum((a - b) ** 2, axis=1))
    
    d = distance(xy)
    

    【讨论】:

    • 此代码有效,但我的数组是 (43,2) 形状。我该怎么做?当我将此代码应用于我的原始数组时,它不起作用。 array([[ 8139, 112607], [ 8139, 115665], [ 8132, 126563], [ 8193, 113938], [ 8193, 123714], [ 8156, 120291], [ 8373, 125253], [ 8400, 131442], [ 8400, 136354], [ 8401, 129352], [ 8439, 129909], [ 8430, 135706], [ 10378, 246013]], dtype=int64)
    • 错误信息是什么?
    • 哦,我修复了它,但我有另一个错误。它说 NameError: name 'int64' is not defined。因为如您所见,我的数组末尾有 dtype=int64 东西。我怎样才能摆脱这个?你能帮我吗:((
    • 删除 dtype=int64 或替换为 dtype=np.int64。 (如果这适合您的需要,请考虑accept my answer)
    • 我怎样才能删除它我是python的新手,很抱歉
    猜你喜欢
    • 2013-11-10
    • 1970-01-01
    • 2013-01-27
    • 1970-01-01
    • 2014-11-29
    • 2018-12-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多