【问题标题】:sklearn GridSearchCV not using sample_weight in score functionsklearn GridSearchCV 未在评分函数中使用 sample_weight
【发布时间】:2018-09-09 21:20:38
【问题描述】:

我有每个样本的权重不同的数据。在我的应用程序中,重要的是在估计模型和比较替代模型时考虑这些权重。

我正在使用sklearn 来估计模型并比较备选的超参数选择。但是这个单元测试显示GridSearchCV 不适用于sample_weights 来估计分数。

有没有办法让sklearn 使用sample_weight 对模型进行评分?

单元测试:

from __future__ import division

import numpy as np
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import log_loss
from sklearn.model_selection import GridSearchCV, RepeatedKFold


def grid_cv(X_in, y_in, w_in, cv, max_features_grid, use_weighting):
  out_results = dict()

  for k in max_features_grid:
    clf = RandomForestClassifier(n_estimators=256,
                                 criterion="entropy",
                                 warm_start=False,
                                 n_jobs=-1,
                                 random_state=RANDOM_STATE,
                                 max_features=k)
    for train_ndx, test_ndx in cv.split(X=X_in, y=y_in):
      X_train = X_in[train_ndx, :]
      y_train = y_in[train_ndx]
      w_train = w_in[train_ndx]
      y_test = y[test_ndx]

      clf.fit(X=X_train, y=y_train, sample_weight=w_train)

      y_hat = clf.predict_proba(X=X_in[test_ndx, :])
      if use_weighting:
        w_test = w_in[test_ndx]
        w_i_sum = w_test.sum()
        score = w_i_sum / w_in.sum() * log_loss(y_true=y_test, y_pred=y_hat, sample_weight=w_test)
      else:
        score = log_loss(y_true=y_test, y_pred=y_hat)

      results = out_results.get(k, [])
      results.append(score)
      out_results.update({k: results})

  for k, v in out_results.items():
    if use_weighting:
      mean_score = sum(v)
    else:
      mean_score = np.mean(v)
    out_results.update({k: mean_score})

  best_score = min(out_results.values())
  best_param = min(out_results, key=out_results.get)
  return best_score, best_param


if __name__ == "__main__":
  RANDOM_STATE = 1337
  X, y = load_iris(return_X_y=True)
  sample_weight = np.array([1 + 100 * (i % 25) for i in range(len(X))])
  # sample_weight = np.array([1 for _ in range(len(X))])

  inner_cv = RepeatedKFold(n_splits=3, n_repeats=1, random_state=RANDOM_STATE)

  outer_cv = RepeatedKFold(n_splits=3, n_repeats=1, random_state=RANDOM_STATE)

  rfc = RandomForestClassifier(n_estimators=256,
                               criterion="entropy",
                               warm_start=False,
                               n_jobs=-1,
                               random_state=RANDOM_STATE)
  search_params = {"max_features": [1, 2, 3, 4]}


  fit_params = {"sample_weight": sample_weight}
  my_scorer = make_scorer(log_loss, 
               greater_is_better=False, 
               needs_proba=True, 
               needs_threshold=False)

  grid_clf = GridSearchCV(estimator=rfc,
                          scoring=my_scorer,
                          cv=inner_cv,
                          param_grid=search_params,
                          refit=True,
                          return_train_score=False,
                          iid=False)  # in this usage, the results are the same for `iid=True` and `iid=False`
  grid_clf.fit(X, y, **fit_params)
  print("This is the best out-of-sample score using GridSearchCV: %.6f." % -grid_clf.best_score_)

  msg = """This is the best out-of-sample score %s weighting using grid_cv: %.6f."""
  score_with_weights, param_with_weights = grid_cv(X_in=X,
                                                   y_in=y,
                                                   w_in=sample_weight,
                                                   cv=inner_cv,
                                                   max_features_grid=search_params.get(
                                                     "max_features"),
                                                   use_weighting=True)
  print(msg % ("WITH", score_with_weights))

  score_without_weights, param_without_weights = grid_cv(X_in=X,
                                                         y_in=y,
                                                         w_in=sample_weight,
                                                         cv=inner_cv,
                                                         max_features_grid=search_params.get(
                                                           "max_features"),
                                                         use_weighting=False)
  print(msg % ("WITHOUT", score_without_weights))

产生输出:

This is the best out-of-sample score using GridSearchCV: 0.135692.
This is the best out-of-sample score WITH weighting using grid_cv: 0.099367.
This is the best out-of-sample score WITHOUT weighting using grid_cv: 0.135692.

解释:由于手动计算不带权重的损失会产生与GridSearchCV 相同的评分,因此我们知道样本权重没有被使用。

【问题讨论】:

    标签: python machine-learning scikit-learn


    【解决方案1】:

    目前在 sklearn 中,GridSearchCV(和任何类继承 BaseSearchCV)仅允许在 **fit_params 中使用 sample_weight,但不能在评分中使用它,这是不正确的,因为 CV 通过未加权选择“最佳估计器”分数。注意,当您 grid.fit(X, y, sample_weight=w) 时,仅在 fit 中使用样本权重,而不是 score。

    有两种方法可以解决这个问题:

    1. 方便的方法:将权重添加为 X 中的第一列。在模型中编写您自定义的评分函数和转换器。
    
    
    from sklearn.base import BaseEstimator, TransformerMixin
    # customized scorer
    def weight_remover_scorer(estimator, X, y):
        y_pred = estimator.predict(X)
        w = X[:,0]
        return your_scorer(y, y_pred, sample_weight=w)
    
    # customized transformer
    class WeightRemover(TransformerMixin, BaseEstimator):
        def fit(self, X, y=None, **fit_params):
            return self
    
        def transform(self, X, y=None, **fit_params):
            return X[:,1:]
    
    # in your main function
    if __name__=='__main__':
        pipe = Pipeline([('remove_weight', WeightRemover()),('model',model)])
        params_grid = {'model__'+k:v for k,v in params_grid.items()}
        X = np.c_[train_w, X]
        X_test = np.c_[test_w, X_test]
        grid = GridSearchCV(pipe, params_grid, cv=5, scoring=weight_remover_scorer)
        grid.fit(X, y)
    
    
    1. 在sklearn 类中添加功能(等待新升级)。只需在BaseSearchCV 中添加参数sample_weight(默认为None),以与fit_params = _check_fit_params(X, fit_params) 相同的方式对它们进行更安全的索引。

    【讨论】:

    • 哦,我忘了在上面的代码中使用样本权重进行拟合。只需将grid.fit(X, y) 更改为grid.fit(X, y, sample_weight=w) 即可!现在我们在 fit 和 score 中使用 sample_weight。
    • 糟糕...由于Pipeline 不支持fit 中的sample_weight,仍然失败。我还尝试在它上面添加一个类包装器,它修改了 fit 函数,如def fit(self, X, y, **fit_params):         w = X[:,0]         self.nX = X.shape[1]-1         return self.wrapped_estimator.fit(X[:,1:], y, sample_weight=w)然而,它是一个本地对象,在尝试并行加速时不能腌制......
    【解决方案2】:

    只是指出正在努力支持这一重要功能:https://github.com/scikit-learn/scikit-learn/pull/13432

    但似乎由于向后兼容性问题以及解决传递任意样本相关信息的更普遍问题的愿望,它花费的时间有点太长了。最后一次尝试似乎是:https://github.com/scikit-learn/scikit-learn/pull/16079

    这里是对该问题的一个很好的评论:http://deaktator.github.io/2019/03/10/the-error-in-the-comparator/

    【讨论】:

      【解决方案3】:

      GridSearchCV 将scoring 作为输入,可以调用。您可以查看如何更改评分功能的详细信息,以及如何通过自己的评分功能here。为了完整起见,以下是该页面中的相关代码:

      EDIT:fit_params 仅传递给 fit 函数,而不传递给 score 函数。如果有应该传递给scorer 的参数,它们应该传递给make_scorer。但这仍然不能解决这里的问题,因为这意味着整个sample_weight 参数将被传递给log_loss,而只有在计算损失时对应于y_test 的部分应该被传递.

      sklearn 不支持这样的事情,但您可以使用padas.DataFrame 破解您的方法。好消息是,sklearn 理解 DataFrame,并保持这种状态。这意味着您可以利用DataFrame 的index,正如您在此处的代码中看到的那样:

        # more code
      
        X, y = load_iris(return_X_y=True)
        index = ['r%d' % x for x in range(len(y))]
        y_frame = pd.DataFrame(y, index=index)
        sample_weight = np.array([1 + 100 * (i % 25) for i in range(len(X))])
        sample_weight_frame = pd.DataFrame(sample_weight, index=index)
      
        # more code
      
        def score_f(y_true, y_pred, sample_weight):
            return log_loss(y_true.values, y_pred,
                            sample_weight=sample_weight.loc[y_true.index.values].values.reshape(-1),
                            normalize=True)
      
        score_params = {"sample_weight": sample_weight_frame}
        my_scorer = make_scorer(score_f,
                                greater_is_better=False, 
                                needs_proba=True, 
                                needs_threshold=False,
                                **score_params)
      
        grid_clf = GridSearchCV(estimator=rfc,
                                scoring=my_scorer,
                                cv=inner_cv,
                                param_grid=search_params,
                                refit=True,
                                return_train_score=False,
                                iid=False)  # in this usage, the results are the same for `iid=True` and `iid=False`
        grid_clf.fit(X, y_frame)
      
        # more code
      

      如您所见,score_f 使用y_true 的index 来查找要使用sample_weight 的哪些部分。为了完整起见,这是整个代码:

      from __future__ import division
      
      import numpy as np
      from sklearn.datasets import load_iris
      from sklearn.ensemble import RandomForestClassifier
      from sklearn.metrics import log_loss
      from sklearn.model_selection import GridSearchCV, RepeatedKFold
      from sklearn.metrics import  make_scorer
      import pandas as pd
      
      def grid_cv(X_in, y_in, w_in, cv, max_features_grid, use_weighting):
        out_results = dict()
      
        for k in max_features_grid:
          clf = RandomForestClassifier(n_estimators=256,
                                       criterion="entropy",
                                       warm_start=False,
                                       n_jobs=1,
                                       random_state=RANDOM_STATE,
                                       max_features=k)
          for train_ndx, test_ndx in cv.split(X=X_in, y=y_in):
            X_train = X_in[train_ndx, :]
            y_train = y_in[train_ndx]
            w_train = w_in[train_ndx]
            y_test = y_in[test_ndx]
      
            clf.fit(X=X_train, y=y_train, sample_weight=w_train)
      
            y_hat = clf.predict_proba(X=X_in[test_ndx, :])
            if use_weighting:
              w_test = w_in[test_ndx]
              w_i_sum = w_test.sum()
              score = w_i_sum / w_in.sum() * log_loss(y_true=y_test, y_pred=y_hat, sample_weight=w_test)
            else:
              score = log_loss(y_true=y_test, y_pred=y_hat)
      
            results = out_results.get(k, [])
            results.append(score)
            out_results.update({k: results})
      
        for k, v in out_results.items():
          if use_weighting:
            mean_score = sum(v)
          else:
            mean_score = np.mean(v)
          out_results.update({k: mean_score})
      
        best_score = min(out_results.values())
        best_param = min(out_results, key=out_results.get)
        return best_score, best_param
      
      
      #if __name__ == "__main__":
      if True:
        RANDOM_STATE = 1337
        X, y = load_iris(return_X_y=True)
        index = ['r%d' % x for x in range(len(y))]
        y_frame = pd.DataFrame(y, index=index)
        sample_weight = np.array([1 + 100 * (i % 25) for i in range(len(X))])
        sample_weight_frame = pd.DataFrame(sample_weight, index=index)
        # sample_weight = np.array([1 for _ in range(len(X))])
      
        inner_cv = RepeatedKFold(n_splits=3, n_repeats=1, random_state=RANDOM_STATE)
      
        outer_cv = RepeatedKFold(n_splits=3, n_repeats=1, random_state=RANDOM_STATE)
      
        rfc = RandomForestClassifier(n_estimators=256,
                                     criterion="entropy",
                                     warm_start=False,
                                     n_jobs=1,
                                     random_state=RANDOM_STATE)
        search_params = {"max_features": [1, 2, 3, 4]}
      
      
        def score_f(y_true, y_pred, sample_weight):
            return log_loss(y_true.values, y_pred,
                            sample_weight=sample_weight.loc[y_true.index.values].values.reshape(-1),
                            normalize=True)
      
        score_params = {"sample_weight": sample_weight_frame}
        my_scorer = make_scorer(score_f,
                                greater_is_better=False, 
                                needs_proba=True, 
                                needs_threshold=False,
                                **score_params)
      
        grid_clf = GridSearchCV(estimator=rfc,
                                scoring=my_scorer,
                                cv=inner_cv,
                                param_grid=search_params,
                                refit=True,
                                return_train_score=False,
                                iid=False)  # in this usage, the results are the same for `iid=True` and `iid=False`
        grid_clf.fit(X, y_frame)
        print("This is the best out-of-sample score using GridSearchCV: %.6f." % -grid_clf.best_score_)
      
        msg = """This is the best out-of-sample score %s weighting using grid_cv: %.6f."""
        score_with_weights, param_with_weights = grid_cv(X_in=X,
                                                         y_in=y,
                                                         w_in=sample_weight,
                                                         cv=inner_cv,
                                                         max_features_grid=search_params.get(
                                                           "max_features"),
                                                         use_weighting=True)
        print(msg % ("WITH", score_with_weights))
      
        score_without_weights, param_without_weights = grid_cv(X_in=X,
                                                               y_in=y,
                                                               w_in=sample_weight,
                                                               cv=inner_cv,
                                                               max_features_grid=search_params.get(
                                                                 "max_features"),
                                                               use_weighting=False)
        print(msg % ("WITHOUT", score_without_weights))
      

      那么代码的输出是:

      This is the best out-of-sample score using GridSearchCV: 0.095439.
      This is the best out-of-sample score WITH weighting using grid_cv: 0.099367.
      This is the best out-of-sample score WITHOUT weighting using grid_cv: 0.135692.
      

      EDIT 2:正如下面的评论所说:

      我的分数与使用此解决方案的 sklearn 分数的差异 起源于我计算加权平均值的方式 分数。如果省略代码的加权平均部分,则两者 输出与机器精度匹配。

      【讨论】:

      • 谢谢,我熟悉make_scorer;这个答案不能解决问题。使用 make_scorer 实现 log_loss 时,我获得了相同的结果(请参阅更新的代码)。值得注意的是,log_loss 函数本身将sample_weights 作为参数;然而,由于输出与加权分数不匹配,我们可以推断fit 没有将sample_weights 传递给my_scorer。换句话说,问题不在于是否使用make_scorer,而在于GridSearchCV 是否能够将样本权重传递给评分函数。
      • 谢谢!这是一个很好的答案。注意——我的分数与使用此解决方案的sklearn 分数的差异源于我计算分数加权平均值的方式。如果省略代码的加权平均部分,则两个输出与机器精度匹配。
      • 非常好的答案,也是唯一一个,我可以在 Stackoverflow 上找到。
      • 有用。可惜sklearn不支持这个:(
      • 这就是我们努力的原因:github.com/scikit-learn/scikit-learn/pull/22083@hipoglucido
      猜你喜欢
      • 2015-10-15
      • 2012-10-14
      • 2018-11-05
      • 2019-03-03
      • 2021-01-09
      • 1970-01-01
      • 2019-08-12
      • 2021-10-03
      • 2020-03-10
      相关资源
      最近更新 更多