【问题标题】:Optimization of a function which look for combination - out-of-memory trouble + speed寻找组合的功能优化 - 内存不足问题+速度
【发布时间】:2013-03-22 17:18:20
【问题描述】:

下面是一个函数,它创建了将 x 的元素分成 n 组的所有可能组合(所有组具有相同数量的元素)

功能:

perm.groups <- function(x,n){
    nx <- length(x)
    ning <- nx/n

    group1 <- 
      rbind(
        matrix(rep(x[1],choose(nx-1,ning-1)),nrow=1),
        combn(x[-1],ning-1)
      )
    ng <- ncol(group1)

    if(n > 2){
      out <- vector('list',ng)

      for(i in seq_len(ng)){
        other <- perm.groups(setdiff(x,group1[,i]),n=n-1)
        out[[i]] <- lapply(seq_along(other),
                       function(j) cbind(group1[,i],other[[j]])
                    )
      }
    out <- unlist(out,recursive=FALSE)
    } else {
      other <- lapply(seq_len(ng),function(i) 
                  matrix(setdiff(x,group1[,i]),ncol=1)
                )
      out <- lapply(seq_len(ng),
                    function(i) cbind(group1[,i],other[[i]])
              )
    }
    out    
}

伪代码(解释)

nb = number of groups
ning = number of elements in every group
if(nb == 2)
   1. take first element, and add it to every possible 
       combination of ning-1 elements of x[-1] 
   2. make the difference for each group defined in step 1 and x 
       to get the related second group
   3. combine the groups from step 2 with the related groups from step 1

if(nb > 2)
   1. take first element, and add it to every possible 
       combination of ning-1 elements of x[-1] 
   2. to define the other groups belonging to the first groups obtained like this, 
       apply the algorithm on the other elements of x, but for nb-1 groups
   3. combine all possible other groups from step 2 
       with the related first groups from step 1

这个函数(和伪代码)最初是由 Joris Meys 在上一篇文章中创建的: Find all possible ways to split a list of elements into a a given number of group of the same size

有没有办法创建一个返回给定数量的随机可能组合的函数? 这样的函数将采用第三个参数,即 percent.possibilities 或 number.possiblities 固定函数返回的随机不同组合的数量。

类似:

new.perm.groups(x=1:12,n=3,number.possiblities=50)

【问题讨论】:

  • 嗯...如果nx 不能被n 整除怎么办?
  • 另外,如果我的计数正确(which, it appears, I have),就会有(nb * ning)! / (nb! * (ning!)^nb) 这样的分区。对于ning 的相对较小的值,这将大于(nb * ning)C(nb)(即二项式系数nb * ning 选择nb),这...是的,那会很大。
  • @JackManey 是的,对不起,我纠正了我的例子。当 nx%%n!=0 时,我收到一条错误消息。
  • 主要问题是存储所有这些分区将占用很多的内存。我不知道您实际上要完成什么,但是如果您可以一次迭代地生成一个分区然后使用它们(不存档它们),那可能会阻止您耗尽内存。
  • @JackManey 是的,您的计算似乎正确! quickyl 的可能性达到非常高的数字!我将有足够的 CPU 来计算东西,但我仍然需要提高计算速度和内存限制

标签: performance r memory out-of-memory


【解决方案1】:

根据@JackManey 的建议,您可以使用等概率方式对一个排列组进行采样

sample.perm.group <- function(ning, ngrp)
{
    if( ngrp==1 ) return(seq_len(ning))

    g1 <- 1+sample(ning*ngrp-1, size=ning-1)

    g1 <- c(1, g1[order(g1)])

    remaining <- seq_len(ning*ngrp)[-g1]

    cbind(g1, matrix(remaining[sample.perm.group(ning, ngrp-1)], nrow=ning), deparse.level=0)
}

其中ning 是每组的元素数,ngrp 是组数。

它返回索引,所以如果你有一个任意向量,你可以将它用作排列:

> ind <- sample.perm.group(3,3)
> ind
     [,1] [,2] [,3]
[1,]    1    2    5
[2,]    3    7    6
[3,]    4    8    9
> LETTERS[1:9][ind]
[1] "A" "C" "D" "B" "G" "H" "E" "F" "I"

要生成大小为 N 的排列样本,您有两种选择:如果您允许重复,即带有替换的样本,您所要做的就是运行前面的函数 N 次。 OTOH,如果您的样本要在没有替换的情况下进行,那么您可以使用拒绝机制:

sample.perm.groups <- function(ning, ngrp, N)
{
    result <- list(sample.perm.group(ning, ngrp))

    for( i in seq_len(N-1) )
    {
        repeat
        {
            y <- sample.perm.group(ning, ngrp)

            if( all(vapply(result, function(x)any(x!=y), logical(1))) ) break
        }

        result[[i+1]] <- y
    }

    result
}

这显然是一种等概率抽样设计,而且不太可能是低效的,因为可能的组合数量通常远大于 N。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2020-09-02
    • 2020-01-26
    • 1970-01-01
    • 1970-01-01
    • 2017-03-12
    • 2019-01-13
    • 2021-08-30
    • 1970-01-01
    相关资源
    最近更新 更多