【问题标题】:Imputation on the test set with fancyimpute用 fancyimpute 对测试集进行插补
【发布时间】:2019-04-18 17:12:16
【问题描述】:

python 包Fancyimpute 提供了几种在 Python 中估算缺失值的方法。该文档提供了以下示例:

# X is the complete data matrix
# X_incomplete has the same values as X except a subset have been replace with NaN

# Model each feature with missing values as a function of other features, and
# use that estimate for imputation.
X_filled_ii = IterativeImputer().fit_transform(X_incomplete)

当将插补方法应用于数据集X 时,这可以正常工作。但是如果training/test 拆分是必要的呢?一次

X_train_filled = IterativeImputer().fit_transform(X_train_incomplete)

被调用,我如何估算测试集并创建X_test_filled?测试集需要使用来自训练集的信息进行估算。我猜IterativeImputer() 应该返回和对象可以适合X_test_incomplete。那可能吗?

请注意,对整个数据集进行插补然后拆分为训练集和测试集是不正确的。

【问题讨论】:

    标签: python missing-data imputation fancyimpute


    【解决方案1】:

    这个包看起来像是在模仿 scikit-learn 的 API。并且在查看源代码后,它看起来确实有一个transform 方法。

    my_imputer = IterativeImputer()
    X_trained_filled = my_imputer.fit_transform(X_train_incomplete)
    
    # now transform test
    X_test_filled = my_imputer.transform(X_test)
    

    插补器将应用它从训练集中学到的相同插补。

    【讨论】:

      猜你喜欢
      • 2020-07-26
      • 2017-12-27
      • 2020-02-01
      • 2019-03-05
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-12-07
      • 2016-11-11
      相关资源
      最近更新 更多