【问题标题】:Given any lat., long. coordinates, what is the fastest way to find the closest coordinates on a list?给定任何纬度,经度。坐标,在列表中找到最近坐标的最快方法是什么?
【发布时间】:2020-09-16 22:38:09
【问题描述】:

我有这样的表格:

import pandas as pd
import numpy as np


df1 = pd.DataFrame([
    ['A', (37.55, 126.97)],
    ['B', (37.56, 126.97)],
    ['C', (37.57, 126.98)]
], columns=['STA_NM', 'COORD'])

df2 = pd.DataFrame([
    ['A-01', (37.57, 126.99)]
], columns=['ID', 'COORD'])

我正在尝试从df2 中选择每个坐标,并从df1 中找到两个最近的站点(STA_NM) 及其到每个坐标的距离,然后将它们添加到df2 的新列中。我尝试了以下代码:

from heapq import nsmallest
from math import cos, asin, sqrt


def dist(x, y):
    p = 0.017453292519943295
    a = 0.5 - cos((y[0] - x[0]) * p) / 2 + cos(x[0] * p) * cos(y[0] * p) * (1 - cos((y[1] - x[1]) * p)) / 2
    return 12741 * asin(sqrt(a))

def shortest(df, v):
    l_sta = []
    
    # get a list of coords
    l_coord = df['COORD'].tolist()
    
    # get the two nearest coordinates
    near_coord = nsmallest(2, l_coord, key=lambda p: dist(v, p))

    # find station names
    l_sta.append((df.loc[df['COORD'] == near_coord[0], 'STA_NM'].to_string(index=False), round(dist(near_coord[0], v) * 1000)))
    l_sta.append((df.loc[df['COORD'] == near_coord[1], 'STA_NM'].to_string(index=False), round(dist(near_coord[1], v) * 1000)))
    
    # e.g.: [('A', 700), ('B', 1000)]
    return l_sta

df2['NEAR_STA'] = df2['COORD'].map(lambda x: shortest(df1, x))

在原始数据中,df1 大约有 700 行,df2 大约有 55k 行。当我尝试上述代码时,花了将近两分钟。有没有更好的方法让它更快?

【问题讨论】:

标签: python python-3.x pandas


【解决方案1】:

在进行距离计算之前,您可以将纬度/经度坐标转换为earth-centered, earth-fixed (ECEF) 坐标(纬度和经度从地核变为 x/y/z)。这将使您的 dist 函数更快,因为它会变成一个单一的欧几里德距离计算。

您也可以放弃 dataframe/lambda 方法并使用 cython 或 numba 来显着加快速度。

如果您知道站点的空间分布情况,也有机会加快速度。例如,如果它们在规则网格上,那么您只需查看四个相邻的站点。如果您知道在某个距离内通常至少有 2 个站点,那么您只需在该半径范围内搜索即可。如果您没有这样的先验信息,那么抱歉没有技巧。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2021-10-28
    • 1970-01-01
    • 2013-01-09
    • 2018-06-18
    • 2020-03-04
    • 1970-01-01
    • 2016-02-03
    相关资源
    最近更新 更多