【问题标题】:Row index for a data.table "binary search" on a subset of columns [duplicate]列子集上的data.table“二进制搜索”的行索引[重复]
【发布时间】:2013-07-15 12:24:54
【问题描述】:

我有一组更大的数据,需要满足特定条件的行数。打包data.table。

days <- strptime(c("2013-01-01 8:00:00", "2013-02-01 8:00:00"), format="%Y-%m-%d %H:%M:%S")
DateTime <- rep(seq(days[1], days[2], length.out=1e6/5), 5)
Update <- rep(LETTERS[3:1], length.out=1e6)
Group <- rep(c("AAA", "BBB", "CCC"), length.out=1e6)
Weight <- trunc(rnorm(1e6, 110, 3))
Weight2 <- rnorm(1e6, 100, 1.5)
DT <- data.table(DateTime, Update, Group, Weight, Weight2)
setkey(DT, DateTime, Update, Group, Weight, Weight2)

Exp <- DT[1e6/2]

我无法创建另一个 data.table 作为没有 DateTime 列的子集,因为该列用于键中。在子集上创建一个新键可能会改变顺序,我需要确定原始顺序被保留。

使用这两个命令可以得到我需要的行号。

system.time(DT[, which(DT$Update==Exp$Update & DT$Group==Exp$Group & DT$Weight==Exp$Weight & DT$Weight2==Exp$Weight2)])
system.time(which(DT$Update==Exp$Update & DT$Group==Exp$Group & DT$Weight==Exp$Weight & DT$Weight2==Exp$Weight2))

但是我需要一种更快的方法来做到这一点。

感谢您的任何建议。

【问题讨论】:

  • 请避免笼统地谈论软件包。他们只会让你的问题变得更长,而且在他们错了的时候特别令人困惑。简单点,我有这个,我试过这个,我明白了,但我想得到这个。
  • 我已经编辑了我的问题。 link 确实为不同但相似的问题提供了答案。解决方案不同。

标签: r data.table


【解决方案1】:

可以通过以下方式获取行号。

which(is.na(DT[list(DT$DateTime, DT$Update, 
DT$Group, DT$Weight, Exp$Weight2), which=TRUE]) == FALSE)

但是它比问题中的向量搜索示例慢 4 倍。

【讨论】:

    猜你喜欢
    • 2020-11-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-08-12
    • 2016-04-20
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多