【问题标题】:All combinations with multiple constraints具有多个约束的所有组合
【发布时间】:2015-10-15 11:31:30
【问题描述】:

我希望生成一组数字的所有可能组合,但有多个约束。我在 Stack Overflow 上发现了几个类似的问题,但似乎都没有解决我的所有限制:

R: sample() command subject to a constraint

R all combinations of 3 vectors with conditions

Generate all combinations given a constraint

R - generate all combinations from 2 vectors given constraints

下面是一个示例数据集。无论如何,在我看来,这是一个确定性数据集。

desired.data <- read.table(text = '
     x1  x2  x3  x4
      1   1   1   1
      1   1   1   2
      1   1   1   3
      1   1   2   1
      1   1   2   2
      1   1   2   3
      1   1   3   3
      1   2   1   1
      1   2   1   2
      1   2   1   3
      1   2   2   1
      1   2   2   2
      1   2   2   3
      1   2   3   3
      1   3   3   3
      0   1   1   1
      0   1   1   2
      0   1   1   3
      0   1   2   1
      0   1   2   2
      0   1   2   3
      0   1   3   3
      0   0   1   1
      0   0   1   2
      0   0   1   3
      0   0   0   1
', header = TRUE, stringsAsFactors = FALSE, na.strings = 'NA')

以下是限制条件:

  1. 第 1 列只能包含 0 或 1
  2. 最后一列只能包含 1、2 或 3
  3. 所有其他列可以包含 0、1、2 或 3
  4. 一旦非 0 出现在一行中,该行的其余部分就不能包含另一个 0
  5. 一旦 3 出现在一行中,则该行的其余部分只能包含 3
  6. 一行中的第一个非 0 数字必须是 1

我知道生成此类数据集的唯一方法是使用嵌套的for-loops,如下所示。我使用这种技术多年,最后决定问是否有更好的方法。

我希望这不是重复的,我希望它不会被认为过于专业。我经常创建这些类型的数据集,一个更简单的解决方案会很有帮助。

my.data <- matrix(0, ncol = 4, nrow = 25)
my.data <- as.data.frame(my.data)

j <- 1

for(i1 in 0:1) {

     if(i1 == 0) i2.begin = 0
     if(i1 == 0) i2.end   = 1
     if(i1 == 1) i2.begin = 1
     if(i1 == 1) i2.end   = 3
     if(i1 == 2) i2.begin = 1
     if(i1 == 2) i2.end   = 3
     if(i1 == 3) i2.begin = 3
     if(i1 == 3) i2.end   = 3

     for(i2 in i2.begin:i2.end) {

          if(i2 == 0) i3.begin = 0
          if(i2 == 0) i3.end   = 1
          if(i2 == 1) i3.begin = 1
          if(i2 == 1) i3.end   = 3
          if(i2 == 2) i3.begin = 1
          if(i2 == 2) i3.end   = 3
          if(i2 == 3) i3.begin = 3
          if(i2 == 3) i3.end   = 3

          for(i3 in i3.begin:i3.end) {

               if(i3 == 0) i4.begin = 1  # 1 not 0 because last column
               if(i3 == 0) i4.end   = 1
               if(i3 == 1) i4.begin = 1
               if(i3 == 1) i4.end   = 3
               if(i3 == 2) i4.begin = 1
               if(i3 == 2) i4.end   = 3
               if(i3 == 3) i4.begin = 3
               if(i3 == 3) i4.end   = 3

               for(i4 in i4.begin:i4.end) {

                    my.data[j,1] <- i1
                    my.data[j,2] <- i2
                    my.data[j,3] <- i3
                    my.data[j,4] <- i4

                    j <- j + 1

               }
          }
     }
}

my.data
dim(my.data)

这是输出:

   V1 V2 V3 V4
1   0  0  0  1
2   0  0  1  1
3   0  0  1  2
4   0  0  1  3
5   0  1  1  1
6   0  1  1  2
7   0  1  1  3
8   0  1  2  1
9   0  1  2  2
10  0  1  2  3
11  0  1  3  3
12  1  1  1  1
13  1  1  1  2
14  1  1  1  3
15  1  1  2  1
16  1  1  2  2
17  1  1  2  3
18  1  1  3  3
19  1  2  1  1
20  1  2  1  2
21  1  2  1  3
22  1  2  2  1
23  1  2  2  2
24  1  2  2  3
25  1  2  3  3
26  1  3  3  3

编辑

抱歉,我最初忘记包含约束 #6。

【问题讨论】:

  • @nongkrong 好点。我需要添加另一个约束:第一个非 0 数字必须是 1。
  • d &lt;- expand.grid(list(0:1,0:2,0:3,1:3)) 这涵盖了 1,2 和 3。4-6,我认为您必须逐行进行。我想不出任何 diff() 或 cumsum() 魔法可以模仿。
  • 如果您不将其转换为 for 循环之前的 data.frame 并像这样一次分配每一行 my.data[j,1]&lt;-c(i1,i2,i3,i4) 在您的示例中,您可以提高 for 循环组的效率'd 必须将矩阵初始化为 26 行而不是 25 行。

标签: r


【解决方案1】:

下面是为此特定示例创建所需数据集的代码。我怀疑代码可以概括。如果我成功概括它,我将发布结果。虽然代码很杂乱而且不直观,但我相信有一个基本的通用模式。

desired.data <- read.table(text = '

     x1  x2  x3  x4
      1   1   1   1
      1   1   1   2
      1   1   1   3
      1   1   2   1
      1   1   2   2
      1   1   2   3
      1   1   3   3
      1   2   1   1
      1   2   1   2
      1   2   1   3
      1   2   2   1
      1   2   2   2
      1   2   2   3
      1   2   3   3
      1   3   3   3
      0   1   1   1
      0   1   1   2
      0   1   1   3
      0   1   2   1
      0   1   2   2
      0   1   2   3
      0   1   3   3
      0   0   1   1
      0   0   1   2
      0   0   1   3
      0   0   0   1
', header = TRUE, stringsAsFactors = FALSE, na.strings = 'NA')

n <- 3   # non-zero numbers
m <- 4-2 # number of middle columns

x1 <- rep(1:0, c(((n*(n-1)) * (n-1) + n), (n*(n-1) + n + (n-1))))
x2 <- rep(c(1:n, 1:0), c(n*m+1, n*m+1, 1, n*m+1, n*1+1))
x3 <- rep(c(rep(1:n, n-1), n, 1:n, 1:0), c(rep(c(n,n,1), n-1), 1, n,n,1, n,1))
x4 <- c(rep(c(rep(1:n, (n-1)), n), (n-1)), n, rep(1:n,(n-1)), n, 1:n, 1)

my.data <- data.frame(x1, x2, x3, x4)

all.equal(desired.data, my.data)
# [1] TRUE

【讨论】:

    【解决方案2】:

    我会使用expand.grid 生成所有组合,然后对其进行子集化,一次一个约束:

    x<-expand.grid(0:1,0:3,0:3,1:3)
    
    ## Once a non-0 appears in a row the rest of that row cannot contain another 0
    b1<-apply(x,1,function(z) min(diff(z!=0))==0)
    x<-x[b1,]
    
    ## Once a 3 appears in a row the rest of that row must only contain 3's
    b1<-apply(x,1,function(z) min(diff(z==3))==0)
    x<-x[b1,]
    
    ## The first non-0 number in a row must be a 1
    b1<-apply(x,1,function(z) {
        w<-which(z==0)
        length(w)==0 || z[tail(w,1)+1]==1
    })
    x<-x[b1,]
    

    现在排序:

    x<-x[order(x[,1],x[,2],x[,3],x[,4]),]
    x
    

    输出:

       Var1 Var2 Var3 Var4
    1     0    0    0    1
    9     0    0    1    1
    41    0    0    1    2
    73    0    0    1    3
    11    0    1    1    1
    43    0    1    1    2
    75    0    1    1    3
    19    0    1    2    1
    51    0    1    2    2
    83    0    1    2    3
    91    0    1    3    3
    12    1    1    1    1
    44    1    1    1    2
    76    1    1    1    3
    20    1    1    2    1
    52    1    1    2    2
    84    1    1    2    3
    92    1    1    3    3
    14    1    2    1    1
    46    1    2    1    2
    78    1    2    1    3
    22    1    2    2    1
    54    1    2    2    2
    86    1    2    2    3
    94    1    2    3    3
    96    1    3    3    3
    

    【讨论】:

    • 谢谢,但当列数可能超过 8 列或要考虑的数字范围为 0:7 而不是 0:3 时,这种方法效果不佳。在这些情况下,我在原始帖子中使用的方法要好得多。
    【解决方案3】:

    类似于@mrip,从expand.grid 开始,它可以处理前3 个约束,因为它们不与其他列交互

    step1<-expand.grid(0:1,0:3,0:3,1:3)
    

    接下来我会过滤它。这种方法与 mrip 的区别在于,我的过滤是一次应用而不是 3 次,因此它应该过滤速度快 3 倍左右。

    filtered<-step1[apply(step1,1,function(x) all(if(length(which(x==0))>0) {max(which(x==0))==length(which(x==0))} else {TRUE}, if(length(which(x==3))>0) {min(which(x==3))==length(x)-length(which(x==3))+1} else {TRUE}, x[!x%in%0][1]==1)),]
    

    应该是这样的。如果你想在这里检查 apply 中的每个元素,它是:

    if(length(which(x==0))&gt;0) {max(which(x==0))==length(which(x==0))} else {TRUE}

    如果有任何零,那么它确保在零之前没有任何内容

    if(length(which(x==3))&gt;0) {min(which(x==3))==length(x)-length(which(x==3))+1} else {TRUE}

    如果有任何 3s,它会确保它们后面没有任何东西。

    x[!x%in%0][1]==1) 这首先从行中过滤出零,然后在该过滤器之后获取行的第一个元素,并且只允许它是一个。

    【讨论】:

    • 我不确定我是否理解。 desired.data 来自哪里?在我的帖子中desired.data 是我希望得到的答案。
    • 变量名称是什么并不重要,这是您问题中变量的唯一名称
    • step1 中正在分析的数据来自哪里?我在问如何在给定一组约束的情况下生成数据集。您似乎在建议如何修改现有数据集。
    • 哦,我错过了您不只是尝试过滤现有数据。我现在在手机上。以后再看吧,修改应该不难。
    • @MarkMiller 没有太大变化,只是从 expand.grid 开始,然后从那里过滤。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-08-22
    • 1970-01-01
    相关资源
    最近更新 更多