【问题标题】:Calculating Standard Error of Coefficients for Logistic Regression in SparkSpark中Logistic回归系数的标准误差计算
【发布时间】:2018-07-07 00:37:37
【问题描述】:

我知道之前有人问过这个问题here。但是我找不到正确的答案。上一篇文章中提供的答案建议使用 Statistics.chiSqTest(data) 提供拟合优度检验(Pearson 卡方检验),而不是用于系数显着性的 Wald 卡方检验。

我试图在 Spark 中为逻辑回归构建参数估计表。我能够获得系数和截距,但我找不到 spark API 来获得系数的标准误差。我看到系数标准误差在线性模型中作为模型摘要的一部分可用。但是逻辑回归模型摘要没有提供这一点。部分示例代码如下。

import org.apache.spark.ml.classification.{BinaryLogisticRegressionSummary, LogisticRegression}

val lr = new LogisticRegression()
  .setMaxIter(10)
  .setRegParam(0.3)
  .setElasticNetParam(0.8)

// Fit the model
val lrModel = lr.fit(training) // Assuming training is my training dataset

val trainingSummary = lrModel.summary
val binarySummary = trainingSummary.asInstanceOf[BinaryLogisticRegressionSummary] // provides the summary information of the fitted model

有什么方法可以计算系数的标准误差。 (或者获取系数的方差-协方差矩阵,从中我们可以得到标准误差)

【问题讨论】:

    标签: apache-spark apache-spark-mllib logistic-regression coefficients standard-error


    【解决方案1】:

    您需要使用带有 Binomial+Logit 的 GLM 方法,而不是 LogisticRegression。

    https://spark.apache.org/docs/2.1.1/ml-classification-regression.html#generalized-linear-regression

    【讨论】:

      猜你喜欢
      • 2018-10-31
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-08-13
      • 2017-12-08
      • 1970-01-01
      • 2014-05-02
      • 2021-09-14
      相关资源
      最近更新 更多