【问题标题】:Quicken this for loop?加快这个 for 循环?
【发布时间】:2014-06-29 06:30:18
【问题描述】:

我有一个名为 cpue 的数据集,有 330 万行。我制作了这个数据框的一个子集,称为 dat.frame。 (请参阅下面的 cpue 和 dat.frame 的负责人。)我在 dat.frame 中添加了两个新字段:“ssh_vec”和“ssh_mag”。虽然 cpue 和 dat.frame 的头部看起来一样,但其余的行实际上并没有相同的顺序。

head(cpue)
  code  event    Lat   Long stat_area Day Month Year id
1  BCO 447602 -43.45 182.73        49  17     3 1995  1



head(dat.frame)
  code  event    Lat   Long stat_area Day Month Year id cal.jdate  ssh_vec  ssh_mag
1  BCO 447602 -43.45 182.73        49  17     3 1995  1   2449857 56.83898 4.499350

目前,我正在运行一个循环,使用唯一标识符“id”将 ssh_vec 和 ssh_mag 变量添加到“cpue”:

cpue$ssh<- NA
cpue$sshmag<- NA

for(i in 1:nrow(dat.frame))
{
    ndx<- dat.frame$id[i]
    cpue_full$ssh[ndx]<- dat.frame$ssh_vec[i]
    cpue_full$sshmag[ndx]<- dat.frame$ssh_mag[i]
}

这已经在周末运行了,只到:

i
[1] 132778

...出自:

nrow(dat.frame)
[1] 2797789

在循环中,没有什么看起来对计算要求太高。有没有更好的选择?

【问题讨论】:

  • 你的sessionInfo()$platform是什么?
  • "x86_64-w64-mingw32/x64(64 位)"

标签: r loops indexing


【解决方案1】:

您确定需要for 循环吗?我认为这可能是等价的:

cpue_full$ssh[dat.frame$id]<- dat.frame$ssh_vec
cpue_full$sshmag[dat.frame$id]<- dat.frame$ssh_mag

【讨论】:

  • 正是我想要的椅子
【解决方案2】:

你还需要循环吗?从发布的代码片段来看,它似乎不是。

cpue_full$ssh[dat.frame$id] <- dat.frame$ssh_vec
cpue_full$sshmag[dat.frame$id]<- dat.frame$ssh_mag

应该可以。一个快速(和小)的虚拟示例:

set.seed(666)
ssh <- rnorm(10^4) 
datf <- data.frame(id = sample.int(10000L), ssh = NA)

system.time(datf$ssh[datf$id] <- ssh) # user 0, system 0, elapsed 0

# Reset dummy data
datf$ssh <- NA 

system.time({
  for (i in 1:nrow(datf) ) {
    ndx <- datf$id[i]
    datf$ssh[ndx] <- ssh[i]
  }
} ) # user 2.26, system 0.02, elapsed 2.28

PS - 我没有使用 data.table 包,所以我不遵循 Ramnath 的回答。一般来说,如果可能的话,你应该避免循环(参见 Fortune(142) 和 The R Inferno 的 Circle 3)。

【讨论】:

  • 看起来我在 nograpes 上面发帖了。对不起,酋长。
  • 谢谢。作为一个循环,在概念上对我来说总是更容易进行计算,我的弱点是使用替代方法来执行它们。谢谢nograpes
【解决方案3】:

我建议您查看data.table。由于我没有你的数据,这里是一个使用虚拟数据的简单示例。

library(data.table)
N = 10^6
dat <- data.table(
  x = rnorm(1000),
  g = sample(LETTERS, N, replace = TRUE)
)

dat2 <- dat[,list(mx = mean(x)),g]

h = merge(dat, dat2, 'g')

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-07-25
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-04-15
    相关资源
    最近更新 更多