【发布时间】:2016-12-17 22:21:47
【问题描述】:
我想创建一个函数,它遍历大量文件,计算每个文件的完整案例数,然后将新行附加到现有数据框中,其中包含文件的“ID”号及其对应的完整病例数。
下面我创建了一个代码,它只返回数据框的最后一行。我相信我的函数只返回最后一行,因为 R 在每个循环中都会覆盖我的数据框,但我不确定。我在网上做了很多研究如何解决这个问题,但我找不到一个简单的解决方案(我对 R 非常陌生)。
您可以在下面看到我的代码和我得到的输出:
complete <- function(directory = "specdata", id = 1:332) {
files_list <- list.files("specdata", full.names = T) # creates a list of files
dat <- data.frame() # creates an emmpty data frame
for (i in id) {
data <- read.csv(files_list[i]) # reads the file "i" in the id vector
nobs <- sum(complete.cases(data)) # counts the number of complete cases in that file
data_frame <- data.frame("ID" = i, nobs) # here I want to store the number of complete cases in a data frame
output <- rbind(dat, data_frame) # here the data_frame should be added to an existing data frame
}
print(output)
}
当我运行complete( , 3:5) 时,我得到以下结果:
ID nobs
1 5 402
感谢四位的帮助! :)
【问题讨论】:
-
我会以不同的方式处理这个问题:(1) 编写一个函数来计算单个文件中完整案例的数量,(2) 使用 lapply() 将该函数应用于文件列表,以及(3) 使用 do.call() 和 rbind() 来构造最终的数据帧。您可以在稍后阶段将所有三个步骤集成到一个函数中。如果没有可重现的示例,编写相应的代码有点困难,所以我将把它作为评论。
标签: r