【问题标题】:Generating random sparse matrix生成随机稀疏矩阵
【发布时间】:2019-10-14 08:17:03
【问题描述】:

我正在处理大型矩阵 - 按10^8 列和10^3-10^4 行的顺序。由于这些矩阵只有一和零(超过 99% 的零),我认为 Matrix 包中的稀疏构造是合适的。但是,我没有看到像下面的示例中那样生成随机矩阵的方法。请注意,非零条目由 column 概率col_prob 定义。

set.seed(1) #For reproducibility
ncols <- 20
nrows <- 10
col_prob <- runif(ncols,0.1,0.2)
rmat <- matrix(rbinom(nrows*ncols,1,col_prob),
       ncol=ncols,byrow=T)

当然,我可以将rmat 转换为稀疏矩阵:

rmat_sparse <- Matrix(rmat, sparse=TRUE)

但是,我想一步生成稀疏矩阵。我不确定Matrix::rsparsematrix这个函数能不能做到这一点。

【问题讨论】:

    标签: r matrix random sparse-matrix


    【解决方案1】:

    以下函数将通过操作空白dgCMatrix 对象的值来生成您正在寻找的类型的稀疏矩阵。它基本上一次创建一个rbinom 行并相应地填充@i 和@p 值。

    library(Matrix)    
    randsparse <- function(nrows, ncols, col_prob) {
      mat <- Matrix(0, nrows, ncols, sparse = TRUE)  #blank matrix for template
      i <- vector(mode = "list", length = ncols)     #each element of i contains the '1' rows
      p <- rep(0, ncols)                             #p will be cumsum no of 1s by column
      for(r in 1:nrows){
        row <- rbinom(ncols, 1, col_prob)            #random row
        p <- p + row                                 #add to column identifier
        if(any(row == 1)){
          for (j in which(row == 1)){
            i[[j]] <- c(i[[j]], r-1)                 #append row identifier
          }
        }
      }
      p <- c(0, cumsum(p))                           #this is the format required
      i <- unlist(i)
      x <- rep(1, length(i))
      mat@i <- as.integer(i)
      mat@p <- as.integer(p)
      mat@x <- x
      return(mat)
    }
    
    set.seed(1)
    randsparse(10, 20, runif(20, 0.1, 0.2))
    
    10 x 20 sparse Matrix of class "dgCMatrix"
    
     [1,] 1 . . . . . . . 1 . . . . . 1 . . . . .
     [2,] . . . . . . . . . . . . . . . . . . . .
     [3,] 1 . . . . . . . . . . . . . . 1 1 . . 1
     [4,] . . . . . . . . . . . . . 1 . . . . . .
     [5,] . . . 1 . . . . 1 . 1 . . . . . . . . .
     [6,] 1 . . . . . . . . . . . . . 1 . . . 1 .
     [7,] . . . . . . . . . . . . . . . . . . . .
     [8,] . 1 . . 1 . . . . . . . 1 . . 1 . . . 1
     [9,] . . 1 . . . . . 1 . . . . 1 . . . 1 . .
    [10,] . . . . . . . . . . 1 . . 1 . . . 1 1 .
    

    【讨论】:

    • 对于大型 nrows=10^4 和 ncols=10^8 来说这会不会很慢?
    • @stats134711 很可能,我没有在这些尺寸上测试过它,但它至少会阻止你在这个过程中使用内存来处理 10^12 元素矩阵!如果你真的想要速度,你可能不得不求助于Rcpp。
    • 是的,我认为这是解决内存+速度问题的唯一方法。
    • @stats134711 您可能会通过重新调整上面的内容以使用sample(c(0,1), replace=TRUE, probs=c(1-p,p)) 生成列来稍微提高速度,我想这可能比rbinom 更快,概率可变。使用上述广泛的过程可能是最简单的,然后在最后转置它,即return(t(mat))
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-06-27
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-08-24
    • 1970-01-01
    相关资源
    最近更新 更多