【发布时间】:2020-07-29 14:52:03
【问题描述】:
我想为数字缺失值估算中位数,为分类缺失值估算模式 然后将所有分类值转换为虚拟变量,居中并缩放它们。 但是,我不想转换客户 ID,也不想将它们居中和缩放。 你能帮我修复我的代码吗?
library(recipes)
train.recipe <- recipe(y ~., data = trainingdata) %>%
step_medianimpute(all_numeric()) %>%
step_modeimpute(all_nominal())
step_dummy(all_nominal(), -all_outcomes(), - trainingdata$Customer_ID) %>%
step_center(all_predictors(), -trainingdata$Customer_ID) %>%
step_scale(all_predictors(), -trainingdata$Customer_ID)
train.recipe %>%
prep() %>%
bake(., data.clean) %>%
glimpse()
【问题讨论】:
-
您能提供一些数据吗?
-
Customer_ID Monthly_Revenue Monthly_Minutes 100072:1 分钟。 : 0.00 分钟。 : 0.0 100403 : 1 中值 : 47.78 中值 : 364.0 100915 : 1 平均值 : 57.92 平均值 : 522.5 (其他):9594 Monthly_Rec_Charge Director_Assisted_Calls Overage_Minutes Min. : 0.00 分钟。 : 0.0000 分钟。 : 0.00 中值 : 45.00 中值 : 0.2500 中值 : 3.00 平均值 : 46.47 平均值 : 0.8954 平均值 : 38.75
-
@TonyFlager 你关心这是如何实现的吗?因为从将客户 ID 转换为行名到简单地使用允许您更明确地命名要转换的列的函数/工作流,有很多答案。那么您需要使用
recipes和step_dummy()的解决方案吗? -
@Fnguyen 不,我不在乎。我只想确保除客户 ID 之外的所有值都已标准化和标准化。你可以帮帮我吗?请给我最简单的方法,因为我是 R 的新手。
标签: r data-modeling prediction missing-data training-data