【问题标题】:GridSearchCV does not report scores on verbose modeGridSearchCV 不报告详细模式的分数
【发布时间】:2021-08-04 03:48:51
【问题描述】:

我在 python 3.8.5 和 sklearn 0.24.1 上使用 GridSearchCV 运行参数网格:

grid_search = GridSearchCV(estimator=xg_clf, scoring=make_scorer(matthews_corrcoef), param_grid=param_grid, n_jobs=args.n_jobs, verbose = 3)

根据文档,

 |  verbose : int
 |      Controls the verbosity: the higher, the more messages.
 |  
 |      - >1 : the computation time for each fold and parameter candidate is
 |        displayed;
 |      - >2 : the score is also displayed;
 |      - >3 : the fold and candidate parameter indexes are also displayed
 |        together with the starting time of the computation.

设置verbose = 3,我做了,应该打印每次运行的马修斯相关系数。

但是,输出是

Fitting 5 folds for each of 480 candidates, totalling 2400 fits
[CV 1/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.9; total time=   0.2s
[CV 2/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.9; total time=   0.2s
[CV 3/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.9; total time=   0.2s
[CV 4/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.9; total time=   0.2s
[CV 5/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.9; total time=   0.2s
[CV 1/5] END colsample_bytree=0.8, gamma=0, learning_rate=0.7, max_depth=3, n_estimators=200, subsample=0.95; total time=   0.2s

为什么GridSearchCV 不为每次运行打印 MCC?

也许这是因为我使用了非标准的记分员?

【问题讨论】:

  • 你在 Google Colab 上工作吗?此外,请确保更新您的库以匹配您引用的文档。
  • @ArturoSbr 我从未听说过 Google Colab。文档来自我笔记本电脑上的命令行。
  • 我提到了 Google Colab,因为它是一个冗长的平台似乎不能很好地工作。无论哪种方式,xg_clf 是 xgboost 对象吗?如果是这样,那可能就是原因。
  • @ArturoSbr xg_clf 确实是一个 XGBoost 对象。 XGBoost 可以与 GridSearchCV 一起使用吗?

标签: python-3.x machine-learning scikit-learn


【解决方案1】:

我用几个不同的 sklearn 版本尝试了类似于你的代码的东西。事实证明,0.24.1 版在verbose=3 时不会打印分数。

这是我使用 sklearn 版本 0.22.2.post1 的代码和输出:

clf = XGBClassifier()
search = GridSearchCV(estimator=clf, scoring=make_scorer(matthews_corrcoef),
                      param_grid={'max_depth':[3, 4, 5]}, verbose=3)
search.fit(X, y)

> Fitting 5 folds for each of 3 candidates, totalling 15 fits
  [CV] max_depth=3 .....................................................
  [CV] ......................... max_depth=3, score=0.959, total=   0.2s

这是我使用 sklearn 版本 0.24.1 的代码和输出:

clf = XGBClassifier()
search = GridSearchCV(estimator=clf, scoring=make_scorer(matthews_corrcoef),
                      param_grid={'max_depth':[3, 4, 5]}, verbose=3)
search.fit(X, y)

> Fitting 5 folds for each of 3 candidates, totalling 15 fits
  [CV 1/5] END ....................................max_depth=3; total time=   0.2s

总之,您发现了一个错误。通常,我建议在 GitHub 上打开一个问题,但您会很高兴知道 0.24.2 版确实打印每个折叠的分数。

您可以尝试pip install scikit-learn --upgrade 或pip install scikit-learn==0.24.1 来解决此问题。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2011-11-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多