【发布时间】:2018-11-01 16:56:08
【问题描述】:
我有一个包含多行的输入数据框。对于每一行,我想应用一个函数。输入数据框有 1,000,000+ 行。如何使用 lapply 加速零件?我想避免使用 Efficient way to apply function to each row of data frame and return list of data frames 中的 apply 系列函数,因为在我的情况下这些方法似乎很慢。
这是一个具有简单功能的可重现示例:
library(tictoc) # enable use of tic() and toc() to record time taken for test to compute
func <- function(coord, a, b, c){
X1 <- as.vector(coord[1])
Y1 <- as.vector(coord[2])
X2 <- as.vector(coord[3])
Y2 <- as.vector(coord[4])
if(c == 0) {
res1 <- mean(c((X1 - a) : (X1 - 1), (Y1 + 1) : (Y1 + 40)))
res2 <- mean(c((X2 - a) : (X2 - 1), (Y2 + 1) : (Y2 + 40)))
res <- matrix(c(res1, res2), ncol=2, nrow=1)
} else {
res1 <- mean(c((X1 - a) : (X1 - 1), (Y1 + 1) : (Y1 + 40)))*b
res2 <- mean(c((X2 - a) : (X2 - 1), (Y2 + 1) : (Y2 + 40)))*b
res <- matrix(c(res1, res2), ncol=2, nrow=1)
}
return(res)
}
## Apply the function
set.seed(1)
n = 10000000
tab <- as.matrix(data.frame(x1 = sample(1:100, n, replace = T), y1 = sample(1:100, n, replace = T), x2 = sample(1:100, n, replace = T), y2 = sample(1:100, n, replace = T)))
tic("test 1")
test <- do.call("rbind", lapply(split(tab, 1:nrow(tab)),
function(x) func(coord = x,
a = 40,
b = 5,
c = 1)))
toc()
## test 1: 453.76 sec elapsed
【问题讨论】:
-
马上想到函数不使用
X2和Y2。 -
另外,如果
coord是data.frame,as.vector(coord[1])和coord[[1]]是一样的,不需要调用函数。 -
其实真正的功能很复杂,我已经简化了。它使用 X1、Y1、X2 和 Y2。此外,函数参数在每个时间步都会改变,但为了简化,我已经删除了循环。因此
tab的值不一样。 -
另一个需要修改的地方是
split(tab, 1:nrow(tab))。这将 df 拆分为ndf,每个只有一行。最好打电话给apply(tab, 1, func)。split单独占用了我的系统。 -
我已经修改了代码,以便函数使用 X2 和 Y2。现在,
coord的值不一样了。