【问题标题】:random forest by group processing/scoring按组处理/评分的随机森林
【发布时间】:2014-04-30 10:05:04
【问题描述】:

我正在尝试使用客户数据库构建预测模型。

我有一个包含 3,000 个客户的数据集。每个客户在测试数据集中有 300 个观察值和 20 个变量(包括因变量)。我还有一个分数数据集,每个唯一的客户 ID 有 50 个观察值和 19 个变量(不包括因变量)。我将测试数据集放在一个单独的文件中,每个客户都由唯一的 ID 变量标识,同样,分数数据集也由唯一的 id 变量标识。

我正在开发一个基于 RandomForest 的预测模型。以下是单个客户的示例。我不确定如何自动为每个客户应用模型并有效地预测和存储模型。

    install.packages(randomForest)
    library(randomForest)
    sales <- read.csv("C:/rdata/test.csv", header=T)
    sales_score <- read.csv("C:/rdata/score.csv", header=T)

  ## RandomForest for Single customer

    sales.rf <- randomForest(Sales ~ ., ntree = 500, data = sales,importance=TRUE)
    sales.rf.test <- predict(sales.rf, sales_score)

我对 SAS 非常熟悉,开始学习 R。对于 SAS 程序员来说,有很多 SAS 程序是通过组处理来实现的,例如:

proc gam data = test;
by id;
model y = x1  x2 x3;
score data = test  out = pred;
run;

此 SAS 程序将为每个唯一 ID 开发一个游戏模型,并将其应用于每个唯一 ID 的测试集。有R等价物吗?

如果有任何例子或想法,我将不胜感激?

非常感谢

【问题讨论】:

  • 您应该需要的唯一“非显而易见”命令是 split。除此之外,这只不过是一个for 循环。只需确保预先分配将包含每个客户的模型和预测值的列表。 (另外,对于大量变量,避免使用randomForest 的公式接口。它通常会慢得多。)
  • 乔丹,谢谢。你能详细说明一下避免公式界面吗?
  • 由于您是 R 新手,因此您应该学习文档。关于公式界面的注释就在注释部分。阅读论据 x 和 y(位于文档顶部)以获取替代方案。

标签: r sas grouping prediction random-forest


【解决方案1】:

假设您的 sales 数据集是 3,000 * 300 = 900,000 行,并且两个数据框都有 customer_id 列,您可以执行以下操作:

pred_groups <- split(seq_len(nrow(sales_score)), sales_score$customer_id)
# pred_groups is now a list, with names the customer_id's and each list
# element an integer vector of row numbers. Now iterate over each customer
# and make predictions on the training set.
preds <- unsplit(structure(lapply(names(pred_groups), function(customer_id) {
  # Train using only observations for this customer.
  # Note we are comparing character to integer but R's natural type
  # coercion should still give the correct answer.
  train_rows <- sales$customer_id == customer_id
  sales.rf <- randomForest(Sales ~ ., ntree = 500,
                           data = sales[train_rows, ],importance=TRUE)

  # Now make predictions only for this customer.
  predict(sales.rf, sales_score[pred_groups[[customer_id]], ])
}), .Names = names(pred_groups)), sales_score$customer_id)

print(head(preds)) # Should now be a vector of predicted scores of length
  # the number of rows in the train set.

编辑:根据@joran,这是一个带有for 的解决方案:

pred_groups <- split(seq_len(nrow(sales_score)), sales_score$customer_id)
preds <- numeric(nrow(sales_score))
for(customer_id in names(pred_groups)) {
  train_rows <- sales$customer_id == customer_id
  sales.rf <- randomForest(Sales ~ ., ntree = 500,
                           data = sales[train_rows, ],importance=TRUE)
  pred_rows <- pred_groups[[customer_id]]
  preds[pred_rows] <- predict(sales.rf, sales_score[pred_rows, ])
})

【讨论】:

  • 没有必要。因为我们在lapply 中,所以使用&lt;- 根本不会修改preds。您的误解是 &lt;&lt;- 是全局变量赋值,因此不鼓励;这是错误的。如果您在父范围内定义变量,它只会修改该变量。将所有这些代码包装在 function() 中不会分配全局 preds 变量。或者,您可以将lapply 替换为for 或使用eval.parent;请注意,assign 不起作用,因为它不会替换作用域。您甚至可以在父环境中使用do.call。无论如何,您的决定是严厉的。
  • 只有在适当的时候才应该遵守规则,但是知道什么时候打破规则很重要。例如,如果没有&lt;&lt;-,实际上不可能对引用类做任何有用的事情!
  • 你对我对
  • @joran 点了。为了完整起见,我添加了一个没有&lt;&lt;- 的版本和一个带有for 循环的版本。
  • 哎呀,关于角色 vs.数字转换,我忘记了,如果您的 customer_id 超过 100,000,这可能是个问题! stackoverflow.com/questions/18964562/…
猜你喜欢
  • 1970-01-01
  • 2015-10-19
  • 1970-01-01
  • 2021-10-28
  • 2018-03-12
  • 1970-01-01
  • 1970-01-01
  • 2019-09-05
  • 2013-09-22
相关资源
最近更新 更多