【问题标题】:Check if each element in a numpy array is in another array检查numpy数组中的每个元素是否在另一个数组中
【发布时间】:2013-04-11 02:34:29
【问题描述】:

这个问题看起来很简单,但我无法找到一个好看的解决方案。我有两个 numpy 数组(A 和 B),我想获取 A 的索引,其中 A 的元素在 B 中,还获取 A 的索引,其中元素不在 B 中。

那么,如果

A = np.array([1,2,3,4,5,6,7])
B = np.array([2,4,6])

目前我正在使用

C = np.searchsorted(A,B)

它利用了A 有序这一事实,并给了我[1, 3, 5],即A 中元素的索引。这很好,但是我如何获得D = [0,2,4,6],A 中不在B 中的元素的索引?

【问题讨论】:

    标签: python numpy


    【解决方案1】:
    import numpy as np
    
    a = np.array([1, 2, 3, 4, 5, 6, 7])
    b = np.array([2, 4, 6])
    c = np.searchsorted(a, b)
    d = np.searchsorted(a, np.setdiff1d(a, b))
    
    d
    #array([0, 2, 4, 6])
    

    【讨论】:

    • 必须搜索两次会减慢速度,最好使用已知的C 来获取D。但是,如果不需要C,这当然是更好的解决方案,所以+1。 (欢迎来到Stack Overflow!)
    【解决方案2】:

    A 中也在 B 中的元素:

    设置(A) & 设置(B)

    A中不在B中的元素:

    集合(A) - 集合(B)

    【讨论】:

    • 这不能回答问题(获取索引,而不是元素)。但是,如果要对 numpy 执行上述操作,请不要将其转换为 set,而是使用 numpy 操作。见intersect1d 和setdiff1d(或最终setxor1d)。
    • 谢谢,因为我在寻找元素而不是索引并且问题标题不明确。我也很欣赏 numpy 的操作。
    【解决方案3】:
    import numpy as np
    
    A = np.array([1,2,3,4,5,6,7])
    B = np.array([2,4,6])
    C = np.searchsorted(A, B)
    
    D = np.delete(np.arange(np.alen(A)), C)
    
    D
    #array([0, 2, 4, 6])
    

    【讨论】:

    • 谢谢!我也喜欢 alexhb 使用 np.setdiff1d 提供的答案。我希望有一个函数可以直接给我索引,但这很好用。
    • 可能有,@Dan,但我想不出。如果您不需要C,请使用他的解决方案,但如果您已经拥有C,我的速度会快一倍。
    【解决方案4】:

    如果不是 B 的每个元素都在 A 中,searchsorted 可能会给你错误的答案。你可以使用numpy.in1d:

    A = np.array([1,2,3,4,5,6,7])
    B = np.array([2,4,6,8])
    mask = np.in1d(A, B)
    print np.where(mask)[0]
    print np.where(~mask)[0]
    

    输出是:

    [1 3 5]
    [0 2 4 6]
    

    但是in1d() 使用排序,这对于大型数据集来说很慢。如果您的数据集很大,您可以使用 pandas:

    import pandas as pd
    np.where(pd.Index(pd.unique(B)).get_indexer(A) >= 0)[0]
    

    时间对比如下:

    A = np.random.randint(0, 1000, 10000)
    B = np.random.randint(0, 1000, 10000)
    
    %timeit np.where(np.in1d(A, B))[0]
    %timeit np.where(pd.Index(pd.unique(B)).get_indexer(A) >= 0)[0]
    

    输出:

    100 loops, best of 3: 2.09 ms per loop
    1000 loops, best of 3: 594 µs per loop
    

    【讨论】:

    • 很高兴知道这种高效的方法,因为我的数据集非常大。非常感谢这个解决方案!
    猜你喜欢
    • 2018-12-17
    • 2021-01-30
    • 2010-10-06
    • 1970-01-01
    • 1970-01-01
    • 2017-03-03
    • 1970-01-01
    • 2021-12-03
    • 1970-01-01
    相关资源
    最近更新 更多