【问题标题】:Indexing takes long time with for loop?for循环索引需要很长时间?
【发布时间】:2016-02-09 08:00:08
【问题描述】:

我正在运行这个for loop 没有任何问题,但这需要很长时间。我想申请家庭可能会更快,但不确定如何。有什么提示吗?

set.seed(1)
nrows <- 1200
ncols <- 1000
outmat <- matrix(NA, nrows, ncols)
dat <- matrix(5, nrows, ncols)
 for (nc in 1 : ncols){
  for(nr in 1 : nrows){
    val <- dat[nr, nc]
    if(!is.na(val)){
      file <- readBin(dir2[val], numeric(), size = 4, n = 1200*1000)
      # my real data where dir2 is a list of files 
      # "dir2 <- list.files("/data/dir2", "*.dat", full.names = TRUE)"
      file <- matrix((data = file), ncol = 1000, nrow = 1200) #my real data

      outmat[nr, nc] <-  file[nr, nc]
    }

  }
}

【问题讨论】:

  • 请您描述一下您的数据。目前尚不清楚为什么不将所有 1200 x 1000 作为单个块读入内存。你有多少这样的积木?我不使用 bin 文件(倾向于使用带有 read.table 或 fread 的 csv 文件)所以可能会漏掉重点。
  • 我不确定您是否有 1200000 个不同的文件,但是您的循环需要很长时间,因为您实际上正在读取 1200000 个文件,并且磁盘访问速度非常慢。使用 apply 不会变得更快。如果您没有太多文件,我建议您将流程还原为首先读取每个文件并存储其数据,然后循环遍历数据以根据需要进行处理。
  • dir2中的几个文件有多少个?

标签: r for-loop


【解决方案1】:

两种解决方案。

如您所说,第一个填充更多内存,但效率更高,如果您有 24 个文件,我想这是可行的。您一次读取所有文件,然后根据dat 正确设置子集。比如:

allContents<-do.call(cbind,lapply(dir2,readBin,n=nrows*ncol,size=4,"numeric")
res<-matrix(allContents[cbind(1:length(dat),c(dat+1))],nrows,ncols)

第二个可以处理稍多的文件(比如 50-100)。它因此读取每个文件和子集的块。您必须打开与获得的文件数量一样多的连接。例如:

outmat <- matrix(NA, nrows, ncols)
connections<-lapply(dir2,file,open="rb")
for (i in 1:ncols)  {
    values<-vapply(connections,readBin,what="numeric",n=nr,size=4,numeric(nr))
    outmat[,i]<-values[cbind(seq_len(nrows),dat[,i]+1)]
}

dat 之后的 +1 是因为,正如您在 cmets 中所述,dat 中的值范围从 0 到 23,R 索引是从 1 开始的。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2010-12-18
    • 1970-01-01
    • 1970-01-01
    • 2012-02-01
    • 1970-01-01
    • 2019-07-21
    • 2019-02-23
    相关资源
    最近更新 更多