【发布时间】:2018-06-12 09:13:11
【问题描述】:
我正在处理一个包含超过 90,000 个 csv.-files 的数据集。每个 csv.-文件显示在特定采样点测量的特定化学品的采样数据。文件如下所示:
#csv1
chemical_ID samplingsite A result year month
1 1 1 0.5 2008 7
1 1 1 0.5 2008 5
1 1 1 0.5 2008 1
1 1 1 0.3 2008 11
1 1 1 0.5 2010 6
1 1 1 0.4 2010 10
1 1 1 0.5 2010 2
1 1 1 0.5 2010 4
1 1 1 0.4 2013 3
1 1 0 0.2 2013 5
1 1 0 0.1 2013 7
1 1 1 0.5 2013 9
1 1 1 0.4 2014 3
1 1 0 0.2 2014 5
1 1 0 0.1 2014 7
1 1 1 0.5 2014 9
#csv2
chemical_ID samplingsite A result year month
2 1 1 0.8 2008 6
2 1 1 0.7 2008 9
2 1 1 0.9 2008 11
2 1 1 0.6 2008 12
2 1 1 0.5 2010 2
2 1 1 0.4 2010 5
2 1 1 0.8 2010 6
2 1 1 0.9 2010 8
#csv3
chemical_ID samplingsite A result year month
100 2 1 1.5 2001 1
100 2 1 1.2 2001 6
100 2 1 1.7 2002 1
100 2 1 0.9 2002 6
100 2 1 1.8 2003 1
100 2 0 1.4 2003 6
100 2 1 1.5 2004 1
100 2 0 1.2 2004 6
为了减少文件数量,我想只选择符合特定条件的文件并将它们保存在新文件夹中。每种化学品的标准应为:
Number of sampled years > 4
Number of samplings per year >= 4
Number of factor “1” in column “A” per year >= 4
我已经尝试过,但找不到我的任务的解决方案,而且谷歌根本没有帮助。这是我到目前为止所得到的:
{
mycsv=list.files(path="D:/…/in ", pattern="allyears")
n <- length(mycsv)
mylist <- vector("list", n)
for(i in 1:n)
mylist[[i]] <- read.csv(mycsv[i], header = TRUE)
mylist <- lapply(mylist, FUN=function(x) length(unique(x$year)))
#???
for(i in 1:n)
write.csv(file = paste("D:/…/out", mycsv[i], sep = ""),
mylist[i], row.names = F)
}
提前致谢
尼斯
【问题讨论】:
-
创建一个函数来读取 csv,检查列是否符合您的条件并返回例如
TRUE如果有,则将所述函数应用于mylist,然后使用结果剔除mylist
标签: r select conditional-statements lapply