【问题标题】:lapply on several subsets of data frame应用于数据框的几个子集
【发布时间】:2018-01-11 09:29:24
【问题描述】:

我有一个数据框 data,其中包含基因组中突变核苷酸的 chromosome 和 position:

structure(list(chrom = c(1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 3L, 
3L, 4L, 4L, 4L, 4L), pos = c(10L, 200L, 134L, 400L, 600L, 1000L, 
20L, 33L, 40L, 45L, 50L, 55L, 100L, 123L)), .Names = c("chrom", 
"pos"), class = "data.frame", row.names = c(NA, -14L))

  chrom  pos
1     1   10
2     1  200
3     1  134
4     1  400
5     1  600
6     1 1000

还有另一个tss_locations,包含gene 和chromosome 中的功能(tss)的位置:

structure(list(gene = structure(c(1L, 4L, 5L, 6L, 7L, 8L, 9L, 
10L, 11L, 2L, 3L), .Label = c("gene1", "gene10", "gene11", "gene2", 
"gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9"
), class = "factor"), chrom = c(1L, 1L, 1L, 2L, 2L, 2L, 3L, 3L, 
3L, 4L, 4L), tss = c(5L, 10L, 23L, 1340L, 313L, 88L, 44L, 57L, 
88L, 74L, 127L)), .Names = c("gene", "chrom", "tss"), class = "data.frame", row.names = c(NA, 
-11L))

   gene chrom  tss
1 gene1     1    5
2 gene2     1   10
3 gene3     1   23
4 gene4     2 1340
5 gene5     2  313
6 gene6     2   88

我正在尝试为data 中的每个pos 计算与在同一条染色体上最近的tss 的距离。

到目前为止,我可以计算每个data$pos 到any tss_locations$tss 的距离(即最接近每个tss 的pos,不考虑染色体):

fun <- function(p) {
  # Get index of nearest tss
  index<-which.min(abs(tss_locations$tss - p))
  # Lookup the value
  closestTss<-tss_locations$tss[[index]]
  # Calculate the distance
  dist<-(closestTss-p)
  list(snp=p, closest=closestTss, distance2nearest=dist)
}

# Run function for each 'pos' in data
dist2tss<-lapply(data$pos, fun)

# Convert to data frame and sort descending:
dist2tss<-do.call(rbind, dist2tss)
dist2tss<-as.data.frame(dist2tss)

dist2tss<-arrange(dist2tss,(as.numeric(distance2nearest)))
dist2tss$distance2nearest<-as.numeric(dist2tss$distance2nearest)

head(dist2tss)

  snp closest distance2nearest
1 600     313             -287
2 400     313              -87
3 200     127              -73
4 100      88              -12
5  33      23              -10
6 134     127               -7

但是,我希望能够为每个pos在同一条染色体上找到最接近的tss。

我知道我可以将它单独应用于每个染色体,但我想查看全局距离(跨所有染色体),但仅在同一染色体上的 positions 和 tsss 之间进行比较。

我该如何调整才能实现这个目标?按染色体对两个数据框进行子集化并合并结果?

到目前为止,这是正确的方法吗?

【问题讨论】:

    标签: r lapply


    【解决方案1】:

    这样的方法可能会为 data 数据帧中的每个染色体获取最接近的 tss。

    data <- structure(list(chrom = c(1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 3L, 
                                     3L, 4L, 4L, 4L, 4L), pos = c(10L, 200L, 134L, 400L, 600L, 1000L, 
                                                                  20L, 33L, 40L, 45L, 50L, 55L, 100L, 123L)), .Names = c("chrom", 
                                                                                                                         "pos"), class = "data.frame", row.names = c(NA, -14L))
    
    tss_locations <- structure(list(gene = structure(c(1L, 4L, 5L, 6L, 7L, 8L, 9L, 
                                                      10L, 11L, 2L, 3L), .Label = c("gene1", "gene10", "gene11", "gene2", 
                                                                                    "gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9"
                                                      ), class = "factor"), chrom = c(1L, 1L, 1L, 2L, 2L, 2L, 3L, 3L, 
                                                                                      3L, 4L, 4L), tss = c(5L, 10L, 23L, 1340L, 313L, 88L, 44L, 57L, 
                                                                                                           88L, 74L, 127L)), .Names = c("gene", "chrom", "tss"), class = "data.frame", row.names = c(NA, 
                                                                                                                                                                                                     -11L))
    
    # Generate needed values by applying function to all rows and transposing t() the results
    data[,c("closest_gene", "closest_tss", "min_dist")] <- t(apply(data, 1, function(x){
       # Get subset of tss_locations where the chromosome matches the current row
       genes <- tss_locations[tss_locations$chrom == x["chrom"], ]
    
       # Find the minimum distance from the current row's pos to the nearest tss location
       min.dist <- min(abs(genes$tss - x["pos"]))
    
       # Find the closest tss location to the current row's pos
       closest_tss <- genes[which.min(abs(genes$tss - x["pos"])), "tss"]
    
       # Check if closest tss location is less than pos and set min.dist to negative if true
       min.dist <- ifelse(closest_tss < x["pos"], min.dist * -1, min.dist)
    
       # Find the closest gene to the current row's pos
       closest_gene <- as.character(genes[which.min(abs(genes$tss - x["pos"])), "gene"])
    
       # Return the values to the matrix
       return(c(closest_gene, closest_tss, min.dist))
    }))
    

    【讨论】:

    • 谢谢 - 看起来很接近,但输出略有偏差。例如。在运行它时,data 的第一行看起来像:1 1 10 gene1 5 0 - 其中最接近的tss 是5,但min_dist 是0
    • 你说得对,我已经编辑了答案来纠正这个问题。
    • 太棒了!我在问题中提供的数据是玩具数据。我的真实数据是一样的,除了data 有一些额外的列。我该如何调整它以适应它(并明确选择pos 列)?
    • data 中的附加列根本不会影响函数。代码已经明确检查了chrom 和pos 列(即x["chrom"] 和x["pos"])
    • 你能在你的答案中添加几行解释吗?我不想要与tss 的绝对距离,但是当我将min.dist &lt;- min(abs(genes$tss - x["pos"])) 更改为min.dist &lt;- min(genes$tss - x["pos"]) 时,min.dist 的值对于负距离是错误的,但对于正距离是正确的 - 为什么会这样?
    猜你喜欢
    • 2019-06-01
    • 1970-01-01
    • 2014-06-08
    • 1970-01-01
    • 2013-11-19
    • 1970-01-01
    • 2014-04-29
    • 1970-01-01
    • 2013-02-10
    相关资源
    最近更新 更多