【问题标题】:Impute Missing Values for Numerical and Categorical, and center and scale the categorical values with exceptions in R估算数值和分类的缺失值,并在 R 中对分类值进行居中和缩放,但有例外
【发布时间】:2020-07-29 14:52:03
【问题描述】:

我想为数字缺失值估算中位数,为分类缺失值估算模式 然后将所有分类值转换为虚拟变量,居中并缩放它们。 但是,我不想转换客户 ID,也不想将它们居中和缩放。 你能帮我修复我的代码吗?

library(recipes)
train.recipe <- recipe(y ~., data = trainingdata) %>%
  step_medianimpute(all_numeric()) %>%
  step_modeimpute(all_nominal())
  step_dummy(all_nominal(), -all_outcomes(), - trainingdata$Customer_ID) %>%
    step_center(all_predictors(), -trainingdata$Customer_ID) %>%
    step_scale(all_predictors(), -trainingdata$Customer_ID)

train.recipe %>%
  prep() %>%
  bake(., data.clean) %>%
  glimpse()

【问题讨论】:

  • 您能提供一些数据吗?
  • Customer_ID Monthly_Revenue Monthly_Minutes 100072:1 分钟。 : 0.00 分钟。 : 0.0 100403 : 1 中值 : 47.78 中值 : 364.0 100915 : 1 平均值 : 57.92 平均值 : 522.5 (其他):9594 Monthly_Rec_Charge Director_Assisted_Calls Overage_Minutes Min. : 0.00 分钟。 : 0.0000 分钟。 : 0.00 中值 : 45.00 中值 : 0.2500 中值 : 3.00 平均值 : 46.47 平均值 : 0.8954 平均值 : 38.75
  • @TonyFlager 你关心这是如何实现的吗?因为从将客户 ID 转换为行名到简单地使用允许您更明确地命名要转换的列的函数/工作流,有很多答案。那么您需要使用recipes 和step_dummy() 的解决方案吗?
  • @Fnguyen 不,我不在乎。我只想确保除客户 ID 之外的所有值都已标准化和标准化。你可以帮帮我吗?请给我最简单的方法,因为我是 R 的新手。

标签: r data-modeling prediction missing-data training-data


【解决方案1】:

在不了解您的数据框并假设客户 ID 是您不想转换的唯一变量的情况下,您可以简单地在转换之前先将 ID 转换为行名:

df %>%
  column_to_rownames("id") # to convert column to rownames

df %>%
  rownames_to_column("id") # to revert that

为此,客户 ID 必须是唯一的!

【讨论】:

    猜你喜欢
    • 2014-10-04
    • 2019-08-10
    • 2014-07-24
    • 2016-02-27
    • 1970-01-01
    • 2019-05-25
    • 2021-09-14
    • 2020-03-27
    • 1970-01-01
    相关资源
    最近更新 更多