【发布时间】:2016-02-09 08:00:08
【问题描述】:
我正在运行这个for loop 没有任何问题,但这需要很长时间。我想申请家庭可能会更快,但不确定如何。有什么提示吗?
set.seed(1)
nrows <- 1200
ncols <- 1000
outmat <- matrix(NA, nrows, ncols)
dat <- matrix(5, nrows, ncols)
for (nc in 1 : ncols){
for(nr in 1 : nrows){
val <- dat[nr, nc]
if(!is.na(val)){
file <- readBin(dir2[val], numeric(), size = 4, n = 1200*1000)
# my real data where dir2 is a list of files
# "dir2 <- list.files("/data/dir2", "*.dat", full.names = TRUE)"
file <- matrix((data = file), ncol = 1000, nrow = 1200) #my real data
outmat[nr, nc] <- file[nr, nc]
}
}
}
【问题讨论】:
-
请您描述一下您的数据。目前尚不清楚为什么不将所有 1200 x 1000 作为单个块读入内存。你有多少这样的积木?我不使用 bin 文件(倾向于使用带有
read.table或fread的 csv 文件)所以可能会漏掉重点。 -
我不确定您是否有 1200000 个不同的文件,但是您的循环需要很长时间,因为您实际上正在读取 1200000 个文件,并且磁盘访问速度非常慢。使用 apply 不会变得更快。如果您没有太多文件,我建议您将流程还原为首先读取每个文件并存储其数据,然后循环遍历数据以根据需要进行处理。
-
dir2中的几个文件有多少个?