【问题标题】:Hyperparameter tuning Random Forest Classifier with GridSearchCV based on probability基于概率的 GridSearchCV 超参数调整随机森林分类器
【发布时间】:2018-02-01 02:48:03
【问题描述】:

刚开始对随机森林二元分类进行超参数调整,我想知道是否有人知道/可以建议如何将评分设置为基于预测概率而不是预测分类。理想情况下,我希望在概率(即 [0.2,0.6,0.7,0.1,0.0])而不是分类(即 [0,1, 1,0,0])。

from sklearn.metrics import roc_auc_score
from sklearn.ensemble import RandomForestClassifier as rfc
from sklearn.grid_search import GridSearchCV

rfbase = rfc(n_jobs = 3, max_features = 'auto', n_estimators = 100, bootstrap=False)

param_grid = {
    'n_estimators': [200,500],
    'max_features': [.5,.7],
    'bootstrap': [False, True],
    'max_depth':[3,6]
}

rf_fit = GridSearchCV(estimator=rfbase, param_grid=param_grid
      , scoring = 'roc_auc')

我认为目前 roc_auc 正在脱离实际分类。在我开始创建自定义评分函数之前,想检查是否有更有效的方法,在此先感谢您的帮助!

【问题讨论】:

  • 查看 skearn 的 make_scorer。它有一个needs_proba 参数。你也许可以从this example 一起摆弄一些东西?
  • 感谢您的快速回复,正是我正在寻找的!如果您想将此移至答案,我会在测试后很高兴地投票:)
  • 确认按预期工作,再次感谢您的快速帮助!
  • 很高兴它有帮助。这取决于您,但您可能想用您发现有效的解决方案来回答您自己的问题。根据 SO,它是 encouraged,尽管某些“专家”不喜欢并驳斥了这一事实。
  • 我在 return roc_auc_score(y_true, y_pred[:,1]) 行中收到“IndexError: too many indices for array”错误

标签: python random-forest hyperparameters


【解决方案1】:

使用 Jarad 提供的参考进行最终求解:

from sklearn.metrics import roc_auc_score
from sklearn.ensemble import RandomForestClassifier as rfc
from sklearn.grid_search import GridSearchCV

rfbase = rfc(n_jobs = 3, max_features = 'auto', n_estimators = 100, bootstrap=False)

param_grid = {
    'n_estimators': [200,500],
    'max_features': [.5,.7],
    'bootstrap': [False, True],
    'max_depth':[3,6]
}

def roc_auc_scorer(y_true, y_pred):
    return roc_auc_score(y_true, y_pred[:, 1])
scorer = make_scorer(roc_auc_scorer, needs_proba=True)

rf_fit = GridSearchCV(estimator=rfbase, param_grid=param_grid
      , scoring = scorer)

【讨论】:

    猜你喜欢
    • 2016-05-11
    • 2017-09-21
    • 2018-03-05
    • 2019-05-01
    • 2018-09-13
    • 2016-01-28
    • 2013-01-10
    • 2014-05-28
    • 2017-04-24
    相关资源
    最近更新 更多