【问题标题】:How to Adjust Classification Threshold with a Spark Decision Tree如何使用 Spark 决策树调整分类阈值
【发布时间】:2017-01-14 12:55:26
【问题描述】:

我正在使用 Spark 2.0 和新的 spark.ml。包。 有没有办法调整分类阈值,以便减少误报的数量。 如果重要的话,我也在使用 CrossValidator。

我看到 RandomForestClassifier 和 DecisionTreeClassifier 都输出一个概率列(我可以手动使用,但 GBTClassifier 没有。

【问题讨论】:

    标签: apache-spark apache-spark-mllib decision-tree


    【解决方案1】:

    听起来您可能正在寻找thresholds 参数:

    final val thresholds: DoubleArrayParam

    Param for Thresholds in multi-class classification 调整概率 预测每个类。数组的长度必须等于 类,值 >= 0。预测具有最大值 p/t 的类, 其中 p 是该类的原始概率, t 是该类' 阈值。

    您需要通过在分类器上调用 setThresholds(value: Array[Double]) 来设置它。

    【讨论】:

    猜你喜欢
    • 2015-10-25
    • 2018-01-12
    • 2013-12-01
    • 2018-01-15
    • 2012-03-18
    • 2016-09-11
    • 2016-07-07
    • 1970-01-01
    • 2017-04-09
    相关资源
    最近更新 更多