【问题标题】:How to apply multiple transforms to the same columns using ColumnTransformer in scikit-learn如何使用 scikit-learn 中的 ColumnTransformer 将多个转换应用于同一列
【发布时间】:2021-10-31 13:42:00
【问题描述】:

我有一个如下所示的数据框:

df = pd.DataFrame(
{
    'x' : range(0,5),
    'y' : [1,2,3,np.nan, np.nan]
})

我想估算 y 的值,并使用以下代码对两个变量进行标准化:

columnPreprocess = ColumnTransformer([
('imputer', SimpleImputer(strategy = 'median'), ['x','y']),   
('scaler', StandardScaler(), ['x','y'])])
columnPreprocess.fit_transform(df)

但是,ColumnTransformer 似乎会为每个步骤设置单独的列,并在不同的列中进行不同的转换。这不是我想要的。

有没有办法对相同的列应用不同的转换并导致输出数组中的列数相同?

【问题讨论】:

    标签: python scikit-learn


    【解决方案1】:

    在这种情况下你应该使用Pipeline:

    import pandas as pd
    import numpy as np
    from sklearn.pipeline import Pipeline
    from sklearn.impute import SimpleImputer
    from sklearn.preprocessing import StandardScaler
    
    df = pd.DataFrame({
        'x': range(0, 5),
        'y': [1, 2, 3, np.nan, np.nan]
    })
    
    pipeline = Pipeline([
        ('imputer', SimpleImputer(strategy='median')),
        ('scaler', StandardScaler())
    ])
    
    pipeline.fit_transform(df)
    # array([[-1.41421356, -1.58113883],
    #        [-0.70710678,  0.        ],
    #        [ 0.        ,  1.58113883],
    #        [ 0.70710678,  0.        ],
    #        [ 1.41421356,  0.        ]])
    

    【讨论】:

      猜你喜欢
      • 2022-08-14
      • 2020-02-03
      • 2022-08-12
      • 2019-04-28
      • 2018-09-19
      • 2015-10-23
      • 2018-05-19
      • 2019-09-05
      • 1970-01-01
      相关资源
      最近更新 更多