【问题标题】:Efficiently merging two data frames on a non-trivial criteria根据重要标准有效地合并两个数据帧
【发布时间】:2013-09-21 08:01:39
【问题描述】:

昨晚回答this question,我花了一个小时试图找到一个没有在for循环中增加data.frame的解决方案,但没有任何成功,所以我很好奇是否有更好的方法关于这个问题。

问题的一般情况归结为:

  • 合并两个data.frames
  • data.frame 中的条目可以在另一个中有 0 个或多个匹配条目。
  • 我们只关心在两者之间有 1 个或多个匹配项的条目。
  • match 函数很复杂,涉及data.frames 中的多个列

对于一个具体的例子,我将使用与链接问题类似的数据:

genes <- data.frame(gene       = letters[1:5], 
                    chromosome = c(2,1,2,1,3),
                    start      = c(100, 100, 500, 350, 321),
                    end        = c(200, 200, 600, 400, 567))
markers <- data.frame(marker = 1:10,
                   chromosome = c(1, 1, 2, 2, 1, 3, 4, 3, 1, 2),
                   position   = c(105, 300, 96, 206, 150, 400, 25, 300, 120, 700))

还有我们复杂的匹配函数:

# matching criteria, applies to a single entry from each data.frame
isMatch <- function(marker, gene) {
  return(
    marker$chromosome == gene$chromosome & 
    marker$postion >= (gene$start - 10) &
    marker$postion <= (gene$end + 10)
  )
}

对于isMatch 为TRUE 的条目,输出应类似于两个data.frame 的sql INNER JOIN。 我尝试构建两个data.frames,以便在另一个data.frame 中可以有0 个或多个匹配项。

我想出的解决方案如下:

joined <- data.frame()
for (i in 1:nrow(genes)) {
   # This repeated subsetting returns the same results as `isMatch` applied across
   # the `markers` data.frame for each entry in `genes`.
   matches <- markers[which(markers$chromosome == genes[i, "chromosome"]),]
   matches <- matches[which(matches$pos >= (genes[i, "start"] - 10)),]
   matches <- matches[which(matches$pos <= (genes[i, "end"] + 10)),]
   # matches may now be 0 or more rows, which we want to repeat the gene for:
   if(nrow(matches) != 0) {
     joined <- rbind(joined, cbind(genes[i,], matches[,c("marker", "position")]))
   }
}

给出结果:

   gene chromosome start end marker position
1     a          2   100 200      3       96
2     a          2   100 200      4      206
3     b          1   100 200      1      105
4     b          1   100 200      5      150
5     b          1   100 200      9      120
51    e          3   321 567      6      400

这是一个相当丑陋和笨拙的解决方案,但我尝试的任何其他方法都失败了:

  • 使用apply,给了我一个list,其中每个元素都是一个矩阵, 无法rbind他们。
  • 我不能先指定joined的尺寸,因为我没有 知道我最终需要多少行。

我相信我以后会想出这个一般形式的问题。那么解决这类问题的正确方法是什么?

【问题讨论】:

  • 当我运行你的代码时,我得到的输出 (joined) 没有任何意义,你能展示你对演示输出的期望吗?
  • 哎呀!有一个错误(&gt;= 在一个实例中应该是&lt;=)。修复并更新了我得到的输出。

标签: r merge data.table


【解决方案1】:

数据表解决方案:滚动连接来满足第一个不等式,然后是矢量扫描来满足第二个不等式。 join-on-first-inequality 将有比最终结果更多的行(因此可能会遇到内存问题),但它会小于 this answer 中的直接合并。

require(data.table)

genes_start <- as.data.table(genes)
## create the start bound as a separate column to join to
genes_start[,`:=`(start_bound = start - 10)]
setkey(genes_start, chromosome, start_bound)

markers <- as.data.table(markers)
setkey(markers, chromosome, position)

new <- genes_start[
    ##join genes to markers
    markers, 
    ##rolling the last key column of genes_start (start_bound) forward
    ##to match the last key column of markers (position)
    roll = Inf, 
    ##inner join
    nomatch = 0
##rolling join leaves positions column from markers
##with the column name from genes_start (start_bound)
##now vector scan to fulfill the other criterion
][start_bound <= end + 10]
##change names and column order to match desired result in question
setnames(new,"start_bound","position")
setcolorder(new,c("chromosome","gene","start","end","marker","position"))
   # chromosome gene start end marker position
# 1:          1    b   100 200      1      105
# 2:          1    b   100 200      9      120
# 3:          1    b   100 200      5      150
# 4:          2    a   100 200      3       96
# 5:          2    a   100 200      4      206
# 6:          3    e   321 567      6      400

可以进行双重连接,但由于它涉及在第二次连接之前重新键入数据表,我认为它不会比上面的矢量扫描解决方案更快。

##makes a copy of the genes object and keys it by end
genes_end <- as.data.table(genes)
genes_end[,`:=`(end_bound = end + 10, start = NULL, end = NULL)]
setkey(genes_end, chromosome, gene, end_bound)

## as before, wrapped in a similar join (but rolling backwards this time)
new_2 <- genes_end[
    setkey(
        genes_start[
        markers, 
        roll = Inf, 
        nomatch = 0
    ], chromosome, gene, start_bound), 
    roll = -Inf, 
    nomatch = 0
]
setnames(new2, "end_bound", "position")

【讨论】:

  • 你绝对赢了。 sqldf 在我的完整数据集上花费了 29 分钟,而 data.table 花费了不到 2 秒!第二个代码块是对的——它的速度几乎是 4 秒以下的两倍。
  • 最终结果是否相同?我只检查了测试数据集,所以我不知道它是否对你的完整数据集有好处。
  • 我没有准确检查,但得出的行数大致相同,所以我假设是这样。
  • 虽然如果我创建类似的start_bound 和end_bound 列,我可能会在sqldf 中提高速度。
  • +1 我还没有完全关注但不能将roll 直接设置为-10 和+10?
【解决方案2】:

我自己通过合并处理了一个非常相似的问题,然后整理出哪些行满足条件。我并不是说这是一个通用的解决方案,如果您正在处理大型数据集,其中匹配条件的条目很少,这可能会效率低下。但要使其适应您的数据:

joined.raw <- merge(genes, markers)
joined <- joined.raw[joined.raw$position >= (joined.raw$start -10) & joined.raw$position <= (joined.raw$end + 10),]
joined
#    chromosome gene start end marker position
# 1           1    b   100 200      1      105
# 2           1    b   100 200      5      150
# 4           1    b   100 200      9      120
# 10          2    a   100 200      4      206
# 11          2    a   100 200      3       96
# 16          3    e   321 567      6      400

【讨论】:

  • 更优雅!但你是对的,创建整个合并的 data.frame 可能并不理想。
  • 是的,我同意 - 这只是根据数据方便
【解决方案3】:

我使用sqldf 包想出的另一个答案。

sqldf("SELECT gene, genes.chromosome, start, end, marker, position 
       FROM genes JOIN markers ON genes.chromosome = markers.chromosome 
       WHERE position >= (start - 10) AND position <= (end + 10)")

使用microbenchmark,它的性能与@alexwhan 的merge 和[ 方法相当。

> microbenchmark(alexwhan, sql)
Unit: nanoseconds
     expr min    lq median  uq  max neval
 alexwhan 435 462.5  468.0 485 2398   100
      sql 422 456.5  466.5 498 1262   100

我还尝试在我拥有的一些相同格式的真实数据上测试这两个函数(genes 为 35,000 行,markers 为 2,000,000 行,joined 输出达到 480,000 行) .

不幸的是,merge 似乎无法处理这么多数据,在joined.raw &lt;- merge(genes, markers) 处出现错误(如果减少行数我不会得到):

Error in merge.data.frame(genes, markers) : 
  negative length vectors are not allowed

sqldf 方法在 29 分钟内成功运行。

【讨论】:

    【解决方案4】:

    在你为我解决了这个问题将近一年之后......现在我花了一些时间通过 awk 使用另一种方式来处理这个......

    awk 'FNR==NR{a[NR]=$0;next}{for (i in a){split(a[i],x," ");if (x[2]==$2 && x[3]-10 <=$3 && x[4]+10 >=$3)print x[1],x[2],x[3],x[4],$0}}' gene.txt makers.txt > genesnp.txt
    

    产生相同结果的那种:

    b   1   100 200 1   1   105
    a   2   100 200 3   2   96
    a   2   100 200 4   2   206
    b   1   100 200 5   1   150
    e   3   321 567 6   3   400
    b   1   100 200 9   1   120
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-01-11
      • 1970-01-01
      • 2022-11-02
      • 1970-01-01
      • 1970-01-01
      • 2017-09-08
      相关资源
      最近更新 更多