【问题标题】:Finding number of elements in one vector that are less than an element in another vector查找一个向量中小于另一个向量中的元素的元素数
【发布时间】:2014-05-21 11:37:09
【问题描述】:

假设我们有几个向量

a <- c(1, 2, 2, 4, 7)
b <- c(1, 2, 3, 5, 7)

对于b 中的每个元素b[i],我想找到a 中小于b[i] 的元素数量,或者,等价的,我想知道c(b[i], a) 中b_i 的等级。

我能想到几种幼稚的方法,例如执行以下任一length(b) 次:

min_rank(c(b[i], a))
sum(a < b[i])

如果length(a) = length(b) = N 其中 N 很大,那么最好的方法是什么?

编辑:

为了澄清,我想知道是否有一种计算效率更高的方法来做到这一点,即在这种情况下我是否可以做得比二次时间更好。

不过,矢量化总是很酷;),谢谢@Henrik!

运行时间

a <- rpois(100000, 20)
b <- rpois(100000, 10)

system.time(
  result1 <- sapply(b, function(x) sum(a < x))
)
# user  system elapsed 
# 71.15    0.00   71.16

sw <- proc.time()
  bu <- sort(unique(b))
  ab <- sort(c(a, bu))
  ind <- match(bu, ab)
  nbelow <- ind - 1:length(bu)
  result2 <- sapply(b, function(x) nbelow[match(x, bu)])
proc.time() - sw

# user  system elapsed 
# 0.46    0.00    0.48 

sw <- proc.time()
  a1 <- sort(a)
  result3 <- findInterval(b - sqrt(.Machine$double.eps), a1)
proc.time() - sw

# user  system elapsed 
# 0.00    0.00    0.03 

identical(result1, result2) && identical(result2, result3)
# [1] TRUE

【问题讨论】:

  • +1 用于提供一个小型、易于重现的玩具数据集,以及您尝试过的代码。为了让希望帮助您检查他们的代码是否产生正确结果的人更容易,请同时发布您想要的输出。干杯。

标签: r sorting vector time-complexity ranking


【解决方案1】:

如果您真的要针对大 N 优化此过程,那么您可能希望至少在开始时删除 b 中的重复值,然后您可以进行排序和匹配:

bu <- sort(unique(b))
ab <- sort(c(a, bu))
ind <- match(bu, ab)
nbelow <- ind - 1:length(bu)

由于我们将 a 和 b 值合并到 ab 中,match 包括所有小于 b 的特定值的 a 以及所有 b,因此我们在最后一行删除了 b 的累积计数。我怀疑这对于大型集合可能更快 - 如果match 内部针对排序列表进行了优化,那么应该是这样,人们希望是这种情况。然后将nbelow 映射回您的原始bs 集应该是一件小事

【讨论】:

  • +1:此解决方案依赖于(快速)排序 - 也就是说,整体复杂度是 n*log(n) 而不是 n²。
【解决方案2】:

假设a越来越弱排序,使用findInterval:

a <- sort(a)
## gives points less than or equal to b[i]
findInterval(b, a)
# [1] 1 3 3 4 5
## to do strictly less than, subtract a small bit from b
## uses .Machine$double.eps (the smallest distinguishable difference)
findInterval(b - sqrt(.Machine$double.eps), a)
# [1] 0 1 3 4 4

【讨论】:

  • 从未有过该功能,非常感谢!希望它在?match的另见部分
  • @GavinKelly 它与pmatch、charmatch 和match.arg 一起位于另请参阅部分。
【解决方案3】:

我并不是说这是“最好的方式”,但它是一种的方式。 sapply 将(匿名)function 应用于b 的每个元素。

 sapply(b, function(x) sum(a < x))
 # [1] 0 1 3 4 4

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-04-19
    • 2019-11-06
    • 1970-01-01
    • 2021-08-03
    • 1970-01-01
    相关资源
    最近更新 更多