【问题标题】:Extract certain values out of a correlation matrix从相关矩阵中提取某些值
【发布时间】:2021-12-16 20:14:42
【问题描述】:

有没有办法将相关系数从相关矩阵中分散出来?

假设我有一个包含 3 个变量(a、b、c)的数据集,我想计算它们之间的相关性。

与


df <- data.frame(a <- c(2, 3, 3, 5, 6, 9, 14, 15, 19, 21, 22, 23),
                 b <- c(23, 24, 24, 23, 17, 28, 38, 34, 35, 39, 41, 43),
                 c <- c(13, 14, 14, 14, 15, 17, 18, 19, 22, 20, 24, 26),
                 d <- c(6, 6, 7, 8, 8, 8, 7, 6, 5, 3, 3, 2))

和

cor(df[, c('a', 'b', 'c')])

我会得到一个相关矩阵:

          a         b         c
 a 1.0000000 0.9279869 0.9604329
 b 0.9279869 1.0000000 0.8942139
 c 0.9604329 0.8942139 1.0000000

有没有办法以这样的方式显示结果:

  1. a 和 b 之间的相关性为:0.9279869。
  2. a 和 c 之间的相关性为:0.9604329。
  3. b 和 c 之间的相关性为:0.8942139:

?

我的相关矩阵明显更大(约 300 个条目),我需要一种方法来仅分散对我重要的值。

谢谢。

【问题讨论】:

    标签: r extract correlation


    【解决方案1】:

    也许你可以试试as.table + as.data.frame

    > as.data.frame(as.table(cor(df[, c("a", "b", "c")])))
      Var1 Var2      Freq
    1    a    a 1.0000000
    2    b    a 0.9279869
    3    c    a 0.9604329
    4    a    b 0.9279869
    5    b    b 1.0000000
    6    c    b 0.8942139
    7    a    c 0.9604329
    8    b    c 0.8942139
    9    c    c 1.0000000
    

    【讨论】:

      【解决方案2】:

      使用 reshape2 和融化

      df <- data.frame("a" = c(2, 3, 3, 5, 6, 9, 14, 15, 19, 21, 22, 23),
                       "b" = c(23, 24, 24, 23, 17, 28, 38, 34, 35, 39, 41, 43),
                       "c" = c(13, 14, 14, 14, 15, 17, 18, 19, 22, 20, 24, 26),
                       "d" = c(6, 6, 7, 8, 8, 8, 7, 6, 5, 3, 3, 2))
      
      tmp=cor(df[, c('a', 'b', 'c')])
      tmp[lower.tri(tmp)]=NA
      diag(tmp)=NA
      
      library(reshape2)
      na.omit(melt(tmp))
      

      导致

        Var1 Var2     value
      4    a    b 0.9279869
      7    a    c 0.9604329
      8    b    c 0.8942139
      

      【讨论】:

        【解决方案3】:

        这是另一种重塑长格式的想法,即

        tidyr::pivot_longer(tibble::rownames_to_column(as.data.frame(cor(df[, c('a', 'b', 'c')])), var = 'rn'), -1)
        
        # A tibble: 9 x 3
          rn    name  value
          <chr> <chr> <dbl>
        1 a     a     1    
        2 a     b     0.928
        3 a     c     0.960
        4 b     a     0.928
        5 b     b     1    
        6 b     c     0.894
        7 c     a     0.960
        8 c     b     0.894
        9 c     c     1    
        

        【讨论】:

        • 谢谢。有没有办法不显示所有引用自己的相关性?就像 a 和 a,b 和 ab = 1?并且只显示单向相关性?只喜欢 a 和 b ,而不是 a 和 b / b 和 a ,因为值是相同的?
        【解决方案4】:

        你可以的,

        df1 = cor(df[, c('a', 'b', 'c')])
        df1 = as.data.frame(as.table(df1))
        df1$Freq = round(df1$Freq,2)
        df2 = subset(df1, (as.character(df1$Var1) != as.character(df1$Var2)))
        df2$res = paste('Correlation between', df2$Var1, 'and', df2$Var2, 'is', df2$Freq)
        
        
         Var1 Var2 Freq                                 res
        2    b    a 0.93 Correlation between b and a is 0.93
        3    c    a 0.96 Correlation between c and a is 0.96
        4    a    b 0.93 Correlation between a and b is 0.93
        6    c    b 0.89 Correlation between c and b is 0.89
        7    a    c 0.96 Correlation between a and c is 0.96
        8    b    c 0.89 Correlation between b and c is 0.89
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2014-04-05
          • 1970-01-01
          • 2018-06-28
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2019-05-09
          相关资源
          最近更新 更多