【问题标题】:Why is pipeline throwing FitFailedWarining when I try LabelEncoder on my pipeline?当我在管道上尝试 LabelEncoder 时,为什么管道会抛出 FitFailedWarning?
【发布时间】:2020-11-28 06:10:12
【问题描述】:

我是机器学习的新手,并试图制作一个项目来让我忙碌,所以我不太了解 sklearn 的工作原理。主要目标是训练模型来预测分类变量。当我尝试 labelEncoding 模型的 y 变量时,我收到以下错误:

ValueError: not enough values to unpack (expected 3, got 2)

  FitFailedWarning)

这是我正在使用的代码

#Rough training

cols_to_use = [col for col in formatData.columns if col not in 'type1']
x = formatData[cols_to_use]
y = formatData.type1
#print(x.columns)
#print(y)


numerical_transformer = SimpleImputer(strategy='constant')
categorical_tansformer = Pipeline(steps=[
                                        ('imputer', SimpleImputer(strategy='most_frequent')),
                                        ('label', LabelEncoder())
                                        ])


preprocessor = ColumnTransformer(transformers=[('num',numerical_transformer),('cat',categorical_tansformer)])

my_pipeline = Pipeline(steps=[('preprocessor',preprocessor),
                              ('model',RandomForestRegressor(n_estimators=50,random_state=0))])

from sklearn.model_selection import cross_validate
from sklearn.model_selection import cross_val_predict

cv_results = cross_validate(my_pipeline,x,y,cv=5,scoring=('r2','neg_mean_absolute_error'))

predictions = cross_val_predict(my_pipeline,x,y,cv=5)
print(cv_results['test_neg_mean_absolute_error'])
print(predictions)

感谢任何帮助,如果您需要更多信息,请发表评论。

【问题讨论】:

    标签: python machine-learning scikit-learn categorical-data label-encoding


    【解决方案1】:

    管道设计用于在X,而不是y。 (围绕此讨论,特别是在重新采样器中,应该在一起改变@ 987654323和y的行进行程;在至少那个方向上看到987654325 @ for修复。)

    特别是,fit_transform(X, y) @ 987654326默认定义为fit(X, y).transform(X)。所以管道中的LabelEncoder 将尝试转换X,并且会失败,因为它不知道如何处理二维输入。您应该只需在管道外标记y 987654330。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-10-03
      • 2010-12-07
      • 2019-07-07
      • 2021-06-23
      • 1970-01-01
      • 2012-04-19
      • 1970-01-01
      • 2022-01-08
      相关资源
      最近更新 更多