【发布时间】:2014-11-21 17:05:15
【问题描述】:
我有一个字符数组,它保存数据框中一行的列名和值。不幸的是,如果特定条目的值为零,则列名和值不会列在数组中。我使用这些信息创建了我想要的数据框,但我依赖于“for 循环”。
我想利用plyr 来避免下面工作代码中的for循环。
types <- c("one", "two", "three") # My data
entry <- c("one(1)", "three(2)") # My data
values <- function(entry, types)
{
frame<- setNames(as.data.frame(matrix(0, ncol = length(types), nrow = 1)), types)
for(s1 in 1:length(entry))
{
name <- gsub("\\(\\w*\\)", "", entry[s1]) # get name
quantity <- as.numeric(unlist(strsplit(entry[s1], "[()]"))[2]) # get value
frame[1, which(colnames(frame)==name)] <- quantity # store
}
return(frame)
}
values(entry, types) # This is how I want the output to look
我尝试了以下方法来拆分数组,但 我不知道如何让 adply 返回单行。
types <- c("one", "two", "three") # data
entry <- c("one(1)", "three(2)") # data
frame<- setNames(as.data.frame(matrix(0, ncol = length(types), nrow = 1)), types)
array_split <- function(entry, frame){
name <- gsub("\\(\\w*\\)", "", entry) # get name
quantity <- as.numeric(unlist(strsplit(entry, "[()]"))[2]) # get value
frame[1, which(colnames(frame)==name)] <- quantity # store
return(frame)
}
adply(entry, 1, array_split, frame)
我应该考虑像 cumsum 这样的东西吗?我想快速完成操作。
【问题讨论】:
-
通常不建议选择“plyr”以提高速度。如果您听说 R 中的循环效率低下,那么您一直在听错顾问。 Hadley 最近一直在开发“dplyr”时考虑到性能,但我认为这不是“plyr”的主要设计目标。在我看来,“plyr”是一种开发统一转换语法的努力。
-
恐怕我经常遇到写得不好的循环,也许使用 plyr 会让我重新思考事情。 cran.r-project.org/doc/Rnews/Rnews_2008-1.pdf 有一个关于高效代码编写的非常好的部分,我将尝试更频繁地记住这些。也感谢您的 dplyr 提示,它看起来很有趣。