【问题标题】:rbind based on columns name and exclude no matchrbind 基于列名并排除不匹配
【发布时间】:2015-04-06 13:53:21
【问题描述】:

样本数据:

l <- list(x=data.frame(X1=1,X2=2,X3=3,X4=4,X5=5),
          y=data.frame(X1=6,X8=7,X4=8,X9=9,X5=10),
          z=data.frame(X1=11,X2=12,X3=13,X4=14,X5=15)
          )

我想rbind这个列表基于预先指定的列名,以便列名(和它的列位置匹配)。

# these are pre-defined columns names we want to `rbind` if no match, exclude the list entry
col <- c("X1","X2","X3","X4","X5") 

所需的输出应该是data.frame:

  X1  X2  X3  X4  X5
   1   2   3   4   5
  11  12  13  14  15

编辑:可能是这样的:

do.call(rbind, lapply(l, function(x) x[!any(is.na(match(c("X1","X2","X3","X4","X5"), names(x))))]))

【问题讨论】:

  • 能有列名c(col, somethingElse)的条目吗?即具有冗余列的条目。
  • @朱利叶斯,不,不可能。必须完全排除不匹配的列表。

标签: r


【解决方案1】:

这是一种方法:

match_all_columns <- function (d, col) {
  if (all(names(d) %in% col)) {
    out <- d[, col]
  } else {
    out <- NULL
  }
  out
}
# or as a one-liner
match_all_columns <- function (d, col) if (all(names(d) %in% col)) d[col]

matched_data <- lapply(l, match_all_columns, col)
result <- do.call(rbind, matched_data)
result
#   X1 X2 X3 X4 X5
# x  1  2  3  4  5
# z 11 12 13 14 15

rbind 知道忽略 NULL 元素。

编辑:我将d[, col] 与d[col] 交换,因为a)它看起来更好,b)如果col 只有一个元素,它可以防止数据帧被丢弃到一个向量中,并且c)我认为它稍微多一点在大型数据帧上表现出色。

【讨论】:

    【解决方案2】:

    还有另一种可能性,允许改变列的顺序:

    output.df <- data.frame(X1=numeric(), X2=numeric(), X3=numeric(),
                            X4=numeric(), X5=numeric())
    
    for(i in seq_along(l)) {
        if(identical(sort(colnames(l[[i]])),sort(colnames(output.df))))
            output.df[nrow(output.df)+1,] <- l[[i]][,colnames(output.df)]
    }
    
    output.df
    
    #   X1 X2 X3 X4 X5
    # 1  1  2  3  4  5
    # 2 11 12 13 14 15
    

    【讨论】:

    • 仅供参考,两种解决方案都已发布忽略列顺序
    • 是的,我实际上只是注意到了这一点。感谢您指出这一点。
    【解决方案3】:

    这似乎也有效:

    do.call(rbind, lapply(l, function(x) x[!any(is.na(match(c("X1","X2","X3","X4","X5"), names(x))))]))
    

    【讨论】:

    • 对我来说似乎是一个非常简单的单行字!你觉得它有什么缺点吗?
    • @DominicComtois,不,我没看到。是的,恐怕它并不适用于所有场合。那么有什么警告呢?
    • 这是一个诚实的问题,我也看不出有什么坏处,只是想知道你为什么最后没有选择那个答案......但你说它并不适用场合?
    • @DominicComtois 我认为这可能比我在具有大量行的数据框中的解决方案慢得多。 R 中的行切片很慢。
    • 哦,我不会接受我的回答!我很高兴看到其他解决方案,我可以从中学习。谢谢。
    【解决方案4】:

    另一个使用data.table的选项

    library(data.table)#v1.9.5+
    na.omit(rbindlist(l, fill=TRUE)[,col, with=FALSE])
    #   X1 X2 X3 X4 X5
    #1:  1  2  3  4  5
    #2: 11 12 13 14 15
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-06-02
      • 2021-05-01
      • 1970-01-01
      • 2012-08-14
      • 2017-06-07
      • 1970-01-01
      相关资源
      最近更新 更多