【发布时间】:2021-10-30 09:55:15
【问题描述】:
我在这里看到了关于“将向量 X 拆分为 R 中的 Y 块”问题的许多变体。例如:here 和 here 仅两个。所以,当我意识到我需要将一个向量分成 Y 个随机大小的块 时,我惊讶地发现随机性要求可能是“新的”——我找不到办法在这里。
所以,这是我拟定的:
k.chunks = function(seq.size, n.chunks) {
break.pts = sample(1:seq.size, n.chunks, replace=F) %>% sort() #Get a set of break points chosen from along the length of the vector without replacement so no duplicate selections.
groups = rep(NA, seq.size) #Set up the empty output vector.
groups[1:break.pts[1]] = 1 #Set the first set of group affiliations because it has a unique start point of 1.
for (i in 2:(n.chunks)) { #For all other chunks...
groups[break.pts[i-1]:break.pts[i]] = i #Set the respective group affiliations
}
groups[break.pts[n.chunks]:seq.size] = n.chunks #Set the last group affiliation because it has a unique endpoint of seq.size.
return(groups)
}
我的问题是:这在某种程度上是不优雅或低效的吗?它会在我计划做的代码中被调用 1000 次,所以效率对我来说很重要。避免for 循环或必须“手动”设置第一组和最后一组会特别好。我的另一个问题:是否有逻辑输入可以打破这一点?我知道n.chunks 不能> seq.size,所以我的意思是除此之外。
【问题讨论】:
-
您需要在这里处理的典型尺寸是多少?对于一个不太大的
n.chunks来说,您的代码应该足够快。 -
另外,您的代码中有两处奇怪的地方。您正在将每个组的最后一个元素重新分配给下一个组 (
break.pts[i-1]:break.pts[i]),最后一个组与之前的分配相同。 -
对于随机 Y,你不会在
sort(sample(1:length(vector), sample(1:length(n.chunks( ie, vector again), replace = FALSE), replace= FALSE))中,除非有最小块大小,此时你可以在seq中折腾? -
@Chris 我认为
sort会缩短你的执行时间。我能想到的最简单的代码是sort(sample(1:n.chunks, seq.size, replace = TRUE))。但这实际上会变得相当缓慢(相对)。 -
@Adam,我同意你的观点并欣赏你下面的代码。我的排序包装只是遵循上面的 OP,但如果不是两者都是我的问题,这是随机的,因为随机化范围的开始/结束似乎是一个非常棘手的问题。
标签: r performance vector vectorization