【问题标题】:How to calculate pValue from GeneralizedLinearRegressionModel using Spark Scala如何使用 Spark Scala 从广义线性回归模型中计算 p 值
【发布时间】:2018-12-28 14:59:14
【问题描述】:

我正在尝试使用 GeneralizedLinearRegression 计算 pValue 并获得以下异常。

    val assembler = new VectorAssembler()
      .setInputCols(final_columns)
      .setOutputCol("Feature")

val glr = new GeneralizedLinearRegression()
      .setFamily("binomial")
      .setLink("logit")
      .setMaxIter(1)
      .setRegParam(0.0)
      .setFeaturesCol("Feature")
      .setLabelCol("LM_2")
      //.setSolver("auto")

    val pipeline = new Pipeline().setStages(Array(assembler,glr))
    val lrModel_general = pipeline.fit(indexedDF)
    val sum = lrModel_general.stages.last.asInstanceOf[GeneralizedLinearRegressionModel].summary.pValues

Exception in thread "main" java.lang.UnsupportedOperationException: No p-value available for this GeneralizedLinearRegressionModel
at org.apache.spark.ml.regression.GeneralizedLinearRegressionTrainingSummary.pValues$lzycompute(GeneralizedLinearRegression.scala:1480)
at org.apache.spark.ml.regression.GeneralizedLinearRegressionTrainingSummary.pValues(GeneralizedLinearRegression.scala:1468)
at com.cvs.scala.ml.model.LR_SqlDB_LocalMessageGrouping$.main(LR_SqlDB_LocalMessageGrouping.scala:172)
at com.cvs.scala.ml.model.LR_SqlDB_LocalMessageGrouping.main(LR_SqlDB_LocalMessageGrouping.scala)

【问题讨论】:

  • 看起来摘要不适用于此模型,当粗麻布不可逆时发生。你应该添加一个检查,我认为它是 model.summary.isAvailable 什么的
  • 当我使用:val sum = lrModel_general.stages.last.asInstanceOf[GeneralizedLinearRegressionModel].hasSummary。它返回 true,表示模型中存在摘要。

标签: scala apache-spark data-science apache-spark-mllib


【解决方案1】:

嗯,首先肯定是关于统计的,所以考虑阅读this answer。

至于您在 Spark 中的解决方案,我建议检查模型的类别并避免对 Ridge 模型进行总结,因为它对这种模型几乎没有用处。

【讨论】:

  • 所以事情是这样的,我使用的数据集在几个列之间具有高度相关性,因此模型无法构建该复杂模式的摘要 cout。在最小化相关性(找出重要特征)后,我能够使用 GLM 摘要获得所有列的 pvalue。干杯!
猜你喜欢
  • 2016-08-24
  • 2010-10-09
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-02-17
  • 2017-12-15
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多