【问题标题】:efficiently look-up data from one data.frame in another data.frame by group有效地从一个数据帧中查找另一个数据帧中的数据。
【发布时间】:2018-12-03 18:06:06
【问题描述】:

我正在为以下问题寻找更快的解决方案。

假设我有以下两个数据集。

df1 <- data.frame(Var1 = c(5011, 2484, 4031, 1143, 7412),
              Var2 = c(2161, 2161, 2161, 2161, 8595))
df2 <- data.frame(team=c("A","A", "B", "B", "B", "C", "C", "D", "D"),
              class=c("5011", "2161", "2484", "4031", "1143", "2161", "5011", "8595", "1143"),
              attribute=c("X1", "X2", "X1", "Z1", "Z2", "Y1", "X1", "Z1", "X2"),
              stringsAsFactors=FALSE)


> df1
  Var1 Var2
1 5011 2161
2 2484 2161
3 4031 2161
4 1143 2161
5 7412 8595

> df2
  team class attribute
1    A  5011        X1
2    A  2161        X2
3    B  2484        X1
4    B  4031        Z1
5    B  1143        Z2
6    C  2161        Y1
7    C  5011        X1
8    D  8595        Z1
9    D  1143        X2

我想知道df2 中的哪些团队在class 中相遇,它们对应于df1 中的。我对行内的顺序不感兴趣。

我当前的代码(粘贴在下面)可以工作,但效率低下。

一些规则:

  • 只有 A 组和 C 组在 df1 中以行形式出现的类中相遇。
  • 团队 B 和团队 D 不会在 df1 中任何成对组合形成一行的班级中相遇。它们被排除在输出之外。

代码:

    teams <- c()
    atts <- c()
    pxs <- unique(df2$team)

    for(j in pxs){
     subs <- subset(df2, team==j)
     for(i in 1:nrow(df1)){
      if(all(df1[i,] %in% subs$class)){
    teams <- rbind(teams, subs$team[i])
    atts <- rbind(atts, subs$attribute)
     } 
     }
    }

    output <- cbind(teams, atts)  

> output
     [,1] [,2] [,3]
[1,] "A"  "X1" "X2"
[2,] "C"  "Y1" "X1"

原始数据由df1df2 中的数百万行组成。

如何更有效地做到这一点?也许通过apply 方法结合data.table

【问题讨论】:

  • 为什么不mergejoin?你的预期输出是什么? merge(df1, df2, by.x = "Var1", by.y = "class")。你能澄清为什么B 不匹配吗?看起来应该。
  • 您能否更清楚地了解您的预期结果?或许提供更多的案例来说明结果中包含和不包含的内容。
  • 2484、4031 和 1143 都出现在 df1 中。你说 B 在 df1 中出现的类中不满足是什么意思?
  • 如果B2484, 9999, 2161 呢?他们是进还是出?
  • - team B does not meet in classes that form a row in df1 - thus not meet the criteria 对于您提供的数据,此编辑仍不清楚。包含/排除规则是什么?

标签: r performance dataframe data.table lookup


【解决方案1】:

不太确定您的规则试图实现什么。

根据您的示例数据、代码和输出,您可能希望先加入 df1 的每一列,然后再内加入 2 个结果:

library(data.table)
setDT(df1)
setDT(df2)[, cls := as.integer(cls)]

#left join df1 with df2 using Var1
v1 <- df2[df1, on=.(cls=Var1)]

#left join df1 with df2 using Var2
v2 <- df2[df1, on=.(cls=Var2)]

#inner join the 2 previous results to ensure that the same team is picked 
#where classes already match in v1 and v2
v1[v2, on=.(team, cls=Var1, Var2=cls), nomatch=0L]

输出:

   team  cls attribute Var2 i.attribute
1:    A 5011        X1 2161          X2
2:    C 5011        X1 2161          Y1

【讨论】:

  • 谢谢!这很好用,比原始代码快得多。
猜你喜欢
  • 2018-01-10
  • 2017-11-18
  • 1970-01-01
  • 2020-10-17
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-01-09
  • 1970-01-01
相关资源
最近更新 更多