【问题标题】:ValueError: Invalid parameter n_estimators for estimator LogisticRegression(random_state=42)ValueError:估计器 LogisticRegression 的参数 n_estimators 无效(random_state=42)
【发布时间】:2021-02-26 09:49:10
【问题描述】:

我已经看过其他类似的问题,但它们对我没有帮助。我正在尝试使用 GridSearchCV。我正在使用三个管道来预测 nfl 播放数据。在网格搜索部分之前它工作得很好。 这是我的代码。

pipe_nfl1_1 = Pipeline([
        ('ssc', StandardScaler()),
        ('lr', LogisticRegression(random_state=42))
])
pipe_nfl1_2 = Pipeline([
        ('mms', MinMaxScaler()),
        ('rfc', RandomForestClassifier(random_state=42))
])
pipe_nfl1_3 = Pipeline([
        ('mms', MinMaxScaler()),
        ('svc', svm.SVC(random_state=42))
])
pipelines1 = [pipe_nfl1_1, pipe_nfl1_2, pipe_nfl1_3]
pipe_dict1 = {0: 'Logistic Regression', 1: 'Random Forest', 2: 'SVC'}
for pipe in pipelines1:
    pipe.fit(X_train1, y_train1)
print('Pipeline test accuracy for predicting 1st downs:')
for idx, val in enumerate(pipelines1):
    print('    %s: %.4f' % (pipe_dict1[idx], val.score(X_test1, y_test1)))
best_acc1 = 0.0
best_clf1 = 0
best_pipe1 = ''
for idx, val in enumerate(pipelines1):
    if val.score(X_test1, y_test1) > best_acc1:
        best_acc1 = val.score(X_test1, y_test1)
        best_pipe1 = val
        best_clf1 = idx
best_acc1 *= 100
print('Classifier with best accuracy for predicting 1st downs is %s with %.2f' % (pipe_dict1[best_clf1], best_acc1) + '%')
param_grid1 = {
        'lr__n_estimators': [2, 4, 6]
}
grid_search1 = GridSearchCV(pipe_nfl1_1, param_grid1, cv=2) 

# fine-tune the hyperparameters
grid_search1.fit(X_train1, y_train1)

# get the best model
final_model1 = grid_search1.best_estimator_
grid_search.best_score_

但我收到一个错误:

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-33-6b0007d9b8f1> in <module>
      2 
      3 # fine-tune the hyperparameters
----> 4 grid_search1.fit(X_train1, y_train1)
      5 
      6 # get the best model

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\utils\validation.py in inner_f(*args, **kwargs)
     70                           FutureWarning)
     71         kwargs.update({k: arg for k, arg in zip(sig.parameters, args)})
---> 72         return f(**kwargs)
     73     return inner_f
     74 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\model_selection\_search.py in fit(self, X, y, groups, **fit_params)
    734                 return results
    735 
--> 736             self._run_search(evaluate_candidates)
    737 
    738         # For multi-metric evaluation, store the best_index_, best_params_ and

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\model_selection\_search.py in _run_search(self, evaluate_candidates)
   1186     def _run_search(self, evaluate_candidates):
   1187         """Search all candidates in param_grid"""
-> 1188         evaluate_candidates(ParameterGrid(self.param_grid))
   1189 
   1190 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\model_selection\_search.py in evaluate_candidates(candidate_params)
    706                               n_splits, n_candidates, n_candidates * n_splits))
    707 
--> 708                 out = parallel(delayed(_fit_and_score)(clone(base_estimator),
    709                                                        X, y,
    710                                                        train=train, test=test,

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\parallel.py in __call__(self, iterable)
   1027             # remaining jobs.
   1028             self._iterating = False
-> 1029             if self.dispatch_one_batch(iterator):
   1030                 self._iterating = self._original_iterator is not None
   1031 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\parallel.py in dispatch_one_batch(self, iterator)
    845                 return False
    846             else:
--> 847                 self._dispatch(tasks)
    848                 return True
    849 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\parallel.py in _dispatch(self, batch)
    763         with self._lock:
    764             job_idx = len(self._jobs)
--> 765             job = self._backend.apply_async(batch, callback=cb)
    766             # A job can complete so quickly than its callback is
    767             # called before we get here, causing self._jobs to

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\_parallel_backends.py in apply_async(self, func, callback)
    206     def apply_async(self, func, callback=None):
    207         """Schedule a func to be run"""
--> 208         result = ImmediateResult(func)
    209         if callback:
    210             callback(result)

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\_parallel_backends.py in __init__(self, batch)
    570         # Don't delay the application, to avoid keeping the input
    571         # arguments in memory
--> 572         self.results = batch()
    573 
    574     def get(self):

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\parallel.py in __call__(self)
    250         # change the default number of processes to -1
    251         with parallel_backend(self._backend, n_jobs=self._n_jobs):
--> 252             return [func(*args, **kwargs)
    253                     for func, args, kwargs in self.items]
    254 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\joblib\parallel.py in <listcomp>(.0)
    250         # change the default number of processes to -1
    251         with parallel_backend(self._backend, n_jobs=self._n_jobs):
--> 252             return [func(*args, **kwargs)
    253                     for func, args, kwargs in self.items]
    254 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\model_selection\_validation.py in _fit_and_score(estimator, X, y, scorer, train, test, verbose, parameters, fit_params, return_train_score, return_parameters, return_n_test_samples, return_times, return_estimator, error_score)
    518             cloned_parameters[k] = clone(v, safe=False)
    519 
--> 520         estimator = estimator.set_params(**cloned_parameters)
    521 
    522     start_time = time.time()

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\pipeline.py in set_params(self, **kwargs)
    139         self
    140         """
--> 141         self._set_params('steps', **kwargs)
    142         return self
    143 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\utils\metaestimators.py in _set_params(self, attr, **params)
     51                 self._replace_estimator(attr, name, params.pop(name))
     52         # 3. Step parameters and other initialisation arguments
---> 53         super().set_params(**params)
     54         return self
     55 

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\base.py in set_params(self, **params)
    259 
    260         for key, sub_params in nested_params.items():
--> 261             valid_params[key].set_params(**sub_params)
    262 
    263         return self

~\AppData\Local\Programs\Python\Python38\lib\site-packages\sklearn\base.py in set_params(self, **params)
    247             key, delim, sub_key = key.partition('__')
    248             if key not in valid_params:
--> 249                 raise ValueError('Invalid parameter %s for estimator %s. '
    250                                  'Check the list of available parameters '
    251                                  'with `estimator.get_params().keys()`.' %

ValueError: Invalid parameter n_estimators for estimator LogisticRegression(random_state=42). Check the list of available parameters with `estimator.get_params().keys()`.

我已经完成 LogisticRegression.get_params().keys() 来获取密钥,但它返回 get_params() 缺少 1 个必需的位置参数:'self'。

【问题讨论】:

  • 更新:我将 __n_estimators 更改为 __C 并且有效

标签: pandas machine-learning logistic-regression hyperparameters gridsearchcv


【解决方案1】:

参数名称中不应包含前导下划线。您希望您的 param_grid1 dict 包含实际上是您正在使用的模型接受的参数的键。那将是 n_estimators 用于 RandomForest,C 用于 LogisticRegression。话虽如此,n_estimators 是模型 RandomForest 的参数,但它不是 LogisticRegression 的参数。 C 是 LogisticRegression 的参数。

我认为您想要做的是对性能最佳的模型的参数空间进行网格搜索,对吧?在这种情况下,您的 param_grid1 变量应更新为性能最佳的模型。您正在测试的模型接受的参数因模型而异。

【讨论】:

    猜你喜欢
    • 2016-12-20
    • 2020-02-22
    • 2018-06-28
    • 2021-08-13
    • 2020-05-19
    • 2020-08-11
    • 2021-05-27
    • 2020-07-12
    • 2016-01-03
    相关资源
    最近更新 更多