【问题标题】:Proper use of "class_weight" parameter in Random Forest classifier在随机森林分类器中正确使用“class_weight”参数
【发布时间】:2020-02-05 01:51:24
【问题描述】:

我有一个多类分类问题,我正在尝试使用随机森林分类器。目标严重不平衡,具有以下分布-

1    34108

4     6748

5     2458

3      132

2       37

7       11

6        6

现在,我正在为 RandomForest 分类器使用“class_weight”参数,据我了解,与类关联的权重采用 {class_label: weight} 的形式

那么,以下是正确的方法吗:

rfc = RandomForestClassifier(n_estimators = 1000, class_weight = {1:0.784, 2: 0.00085, 3: 0.003, 4: 0.155, 5: 0.0566, 6: 0.00013, 7: 0.000252})

感谢您的帮助!

【问题讨论】:

标签: machine-learning scikit-learn classification random-forest


【解决方案1】:

如果您选择class_weight = "balanced",则类的权重将与它们在数据中出现的频率成反比。

在您的示例中,您对代表人数过多的班级的权重比对代表人数不足的班级的权重更大。我相信这与您想要实现的目标相反。

计算每个类权重的基本公式是total observations / (number of classes * observations in class)。

【讨论】:

    猜你喜欢
    • 2018-05-20
    • 2015-08-28
    • 2018-04-10
    • 2018-05-27
    • 2018-02-18
    • 2020-04-27
    • 2017-09-10
    • 2018-03-05
    • 2020-07-02
    相关资源
    最近更新 更多