【问题标题】:Caret confusionMatrix measures are wrong?插入符号混淆矩阵的措施是错误的?
【发布时间】:2020-10-28 20:39:21
【问题描述】:

我做了一个函数来计算混淆矩阵的敏感性和特异性,后来才发现caret 包有一个,confusionMatrix()。当我尝试它时,事情变得非常混乱,因为它似乎 caret 使用了错误的公式??

示例数据:

dat <- data.frame(real = as.factor(c(1,1,1,0,0,1,1,1,1)),
                  pred = as.factor(c(1,1,0,1,0,1,1,1,0)))
cm <- table(dat$real, dat$pred)
cm
    0 1
  0 1 1
  1 2 5

我的功能:

model_metrics <- function(cm){
  acc <- (cm[1] + cm[4]) / sum(cm[1:4])
  # accuracy = ratio of the correctly labeled subjects to the whole pool of subjects = (TP+TN)/(TP+FP+FN+TN)
  sens <- cm[4] / (cm[4] + cm[3])
  # sensitivity/recall = ratio of the correctly +ve labeled to all who are +ve in reality = TP/(TP+FN)
  spec <- cm[1] / (cm[1] + cm[2])
  # specificity = ratio of the correctly -ve labeled cases to all who are -ve in reality = TN/(TN+FP)
  err <- (cm[2] + cm[3]) / sum(cm[1:4]) #(all incorrect / all)
  metrics <- data.frame(Accuracy = acc, Sensitivity = sens, Specificity = spec, Error = err)
  return(metrics)
}

现在将confusionMatrix() 的结果与我的函数的结果进行比较:

library(caret)
c_cm <- confusionMatrix(dat$real, dat$pred)
c_cm
          Reference
Prediction 0 1
         0 1 1
         1 2 5
c_cm$byClass
Sensitivity          Specificity       Pos Pred Value       Neg Pred Value            Precision               Recall 
  0.3333333            0.8333333            0.5000000            0.7142857            0.5000000            0.3333333

model_metrics(cm)
  Accuracy Sensitivity Specificity     Error
1 0.6666667   0.8333333   0.3333333 0.3333333

敏感性和特异性似乎在我的函数和confusionMatrix() 之间互换。我以为我使用了错误的公式,但我仔细检查了Wiki,我是对的。我还仔细检查了我是否从混淆矩阵中调用了正确的值,我很确定我是。 caret documentation 也表明它使用了正确的公式,所以我不知道发生了什么。

caret 函数是否错误,或者(更有可能)我犯了一些令人尴尬的明显错误?

【问题讨论】:

    标签: r machine-learning r-caret confusion-matrix


    【解决方案1】:

    插入符号没有错。

    首先。考虑如何构建表格。 table(first, second) 将生成一个表,其中行中包含 first,列中包含 second。

    此外,在对表格进行子集化时,应按列计算单元格。例如,在您的函数中,计算灵敏度的正确方法是

     sens <- cm[4] / (cm[4] + cm[2])
    

    最后,阅读一个函数的帮助页面总是一个好主意,它没有给你预期的结果。 ?confusionMatrix 会给你帮助页面。

    在对这个函数执行此操作时,您会发现您可以指定将哪个因子水平视为肯定结果(使用positive 参数)。

    另外,请注意如何使用该功能。为避免混淆,我建议使用命名参数而不是按位置依赖参数规范。

    第一个参数是数据(预测类别的一个因素),第二个参数引用是观察类别的一个因素(在您的情况下为dat$real)。

    要得到你想要的结果:

    confusionMatrix(data = dat$pred, reference = dat$real, positive = "1")
    
    Confusion Matrix and Statistics
    
              Reference
    Prediction 0 1
             0 1 2
             1 1 5
                                              
                   Accuracy : 0.6667          
                     95% CI : (0.2993, 0.9251)
        No Information Rate : 0.7778          
        P-Value [Acc > NIR] : 0.8822          
                                              
                      Kappa : 0.1818          
                                              
     Mcnemar's Test P-Value : 1.0000          
                                              
                Sensitivity : 0.7143          
                Specificity : 0.5000          
             Pos Pred Value : 0.8333          
             Neg Pred Value : 0.3333          
                 Prevalence : 0.7778          
             Detection Rate : 0.5556          
       Detection Prevalence : 0.6667          
          Balanced Accuracy : 0.6071          
                                              
           'Positive' Class : 1 
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-03-20
      • 2021-04-16
      • 1970-01-01
      • 2018-09-07
      • 2013-11-21
      • 2019-10-15
      • 2015-09-18
      • 1970-01-01
      相关资源
      最近更新 更多