【问题标题】:R select multiple files out of a list of files by criteria in columns and save in a new folderR按列中的条件从文件列表中选择多个文件并保存在新文件夹中
【发布时间】:2018-06-12 09:13:11
【问题描述】:

我正在处理一个包含超过 90,000 个 csv.-files 的数据集。每个 csv.-文件显示在特定采样点测量的特定化学品的采样数据。文件如下所示:

#csv1
chemical_ID samplingsite    A   result  year    month   
1   1   1   0.5 2008    7
1   1   1   0.5 2008    5
1   1   1   0.5 2008    1
1   1   1   0.3 2008    11
1   1   1   0.5 2010    6
1   1   1   0.4 2010    10
1   1   1   0.5 2010    2
1   1   1   0.5 2010    4
1   1   1   0.4 2013    3
1   1   0   0.2 2013    5
1   1   0   0.1 2013    7
1   1   1   0.5 2013    9
1   1   1   0.4 2014    3
1   1   0   0.2 2014    5
1   1   0   0.1 2014    7
1   1   1   0.5 2014    9

#csv2
chemical_ID samplingsite    A   result  year    month
2   1   1   0.8 2008    6
2   1   1   0.7 2008    9
2   1   1   0.9 2008    11
2   1   1   0.6 2008    12
2   1   1   0.5 2010    2
2   1   1   0.4 2010    5
2   1   1   0.8 2010    6
2   1   1   0.9 2010    8

#csv3
chemical_ID samplingsite    A   result  year    month
100 2   1   1.5 2001    1
100 2   1   1.2 2001    6
100 2   1   1.7 2002    1
100 2   1   0.9 2002    6
100 2   1   1.8 2003    1
100 2   0   1.4 2003    6
100 2   1   1.5 2004    1
100 2   0   1.2 2004    6

为了减少文件数量,我想只选择符合特定条件的文件并将它们保存在新文件夹中。每种化学品的标准应为:

Number of sampled years > 4
Number of samplings per year >= 4
Number of factor “1” in column “A” per year >= 4

我已经尝试过,但找不到我的任务的解决方案,而且谷歌根本没有帮助。这是我到目前为止所得到的:

{
mycsv=list.files(path="D:/…/in ", pattern="allyears")
n <- length(mycsv)
mylist <- vector("list", n)

for(i in 1:n)
mylist[[i]] <- read.csv(mycsv[i], header = TRUE)
mylist <- lapply(mylist, FUN=function(x) length(unique(x$year)))
#???

for(i in 1:n)
write.csv(file = paste("D:/…/out", mycsv[i], sep = ""), 
mylist[i], row.names = F)
}

提前致谢
尼斯

【问题讨论】:

  • 创建一个函数来读取 csv,检查列是否符合您的条件并返回例如TRUE 如果有,则将所述函数应用于mylist,然后使用结果剔除mylist

标签: r select conditional-statements lapply


【解决方案1】:

此方法获取您的文件列表,创建一个函数来检查您的条件,使用该函数检查文件是否符合这些条件,创建一个符合您的函数的文件列表,然后将 csvs 写入一个新文件夹(您必须已创建)。

该示例旨在与您问题中的 csvs 一起使用,当我解释它们时,没有一个符合您的标准,因此在 csv1 符合您的标准的地方添加了一个测试标准。要关闭这些,只需从您的标准中删除#,然后在测试标准前添加#。

file.list <- list.files() # gets list of files - assumes your working directory is where the files are

check.csv <- function(csv.path){ #checks your criteria

  the.csv <- read.csv(file = csv.path, header = TRUE)

  sampled.years <- length(unique(the.csv$year))

  min.samples.per.year <- min(table(the.csv$year))

  min.f1A <- min(table(the.csv$year, the.csv$A)[,"1"])

  #your criteria
  #meets.criteria <- ifelse(sampled.years > 4 & min.samples.per.year >= 4 & min.f1A >=4, TRUE, FALSE) 

  #test criteria
  meets.criteria <- ifelse(sampled.years >= 4 & min.samples.per.year >= 4 & min.f1A >= 2, TRUE, FALSE)

  return(meets.criteria)  
} 

check.files <- sapply(file.list, check.csv) # checks if files meet criteria, as above, assumes that file.list has the whole path, which it will if your working directory is where the files are

files.to.write <- file.list[check.files] # subsets list of files to move

read.write <- function(csv.path){ # function to write csvs into new folder specified in the path as other_folder

  the.csv <- read.csv(file = csv.path, header = TRUE)

  write.csv(the.csv, file = sprintf("other_folder/%s", csv.path))
  # this other_folder must exist
} 


sapply(files.to.write, read.write) # write csvs to new folder

【讨论】:

  • 您好,非常感谢。这段代码应该正是我正在寻找的。不幸的是,我在代码末尾收到两条错误消息:文件中的错误(文件,ifelse(附加,“a”,“w”)):无法打开连接,此外:警告消息:在文件中(文件, ifelse(append, "a", "w")) : 无法打开文件 'D:/_..._final/': 权限被拒绝——我错过了什么?
  • 这些错误发生在哪里?在sapply(files.to.write, read.write) 通话之后?这可能与计算机上的管理员权限有关?你能write.csv 像cars 数据这样的任意对象吗?
  • 是的,它们发生在 sapply(files.to.write, read.write) 调用之后。通常我对 write.csv 没有任何问题,我是这台笔记本电脑的管理员。
  • 你能运行read.write(files.to.write[1])吗?如果没有,请尝试函数外部的每个命令。不清楚问题出在此代码还是特定于您的机器的其他问题上。您说过您“通常”对 write.csv 没有任何问题,但是您可以编写其中一个文件吗?
  • 我可以通过将write.csv(the.csv, file = sprintf("D:/.../out", csv.path)) 更改为write.csv(the.csv, file = paste("D:/.../out/new", csv.path, sep = "")) 来解决问题。现在代码完美运行。
【解决方案2】:

首先创建文件夹的副本。然后我们将删除所有不满足上述条件的文件:

写一个函数:

testfile = function(path){
dat= read.csv(path,header=T)
dat1=aggregate(.~year,dat,length)
a=nrow(d)>4
b=d$samplingsite>=4
d=d$A>-4
if(!(a&b&d)) file.remove(path)
}

mycsv=list.files(path="D:/…/in ", pattern="allyears")# Path to the duplicate folder

sapply(mycsv,testfile)

【讨论】:

    猜你喜欢
    • 2017-11-29
    • 1970-01-01
    • 2018-08-20
    • 2019-06-18
    • 2021-01-28
    • 1970-01-01
    • 1970-01-01
    • 2021-01-08
    • 1970-01-01
    相关资源
    最近更新 更多