【问题标题】:Implementing a function in R to find the coefficients of a linear regression model from the ground up在 R 中实现一个函数以从头开始查找线性回归模型的系数
【发布时间】:2020-01-26 02:38:08
【问题描述】:
# but cannot handle categorical variables
my_lm <- function(explanatory_matrix, response_vec) {
  exp_mat <- as.matrix(explanatory_matrix)
  intercept <- rep(1, nrow(exp_mat))
  exp_mat <- cbind(exp_mat, intercept)
  solve(t(exp_mat) %*% exp_mat) %*% (t(exp_mat) %*% response_vec)
}

当 explanatory_matrix 中有分类变量时,上述代码将不起作用。 我该如何实现?

【问题讨论】:

  • 使用 one-hot 编码将它们转换为数值 (discussion)
  • 见?model.matrix ...
  • @athletic_coder 我的解决方案对你有用吗?

标签: r linear-regression lm


【解决方案1】:

这是一个具有一个分类变量的数据集的示例:

set.seed(123)

x <- 1:10
a <- 2
b <- 3

y <- a*x + b + rnorm(10)
# categorical variable
x2 <- sample(c("A", "B"), 10, replace = T)
# one-hot encoding
x2 <- as.integer(d$x2 == "A")


xm <- matrix(c(x, x2, rep(1, length(x))), ncol = 3, nrow = 10)
ym <- matrix(y, ncol = 1, nrow = 10)

beta_hat <- MASS::ginv(t(xm) %*% xm) %*% t(xm) %*% ym
beta_hat

这给出(注意系数的顺序 - 它与预测列的顺序相匹配):

      [,1]
[1,]  1.9916754
[2,] -0.7594809
[3,]  3.2723071

与lm的输出相同:

d <- data.frame(x = x,
                x2 = x2,
                y = y)
lm(y ~ ., data = d)

输出

# Call:
#   lm(formula = y ~ ., data = d)
# 
# Coefficients:
#   (Intercept)            x           x2  
# 3.2723       1.9917      -0.7595  

【讨论】:

    【解决方案2】:

    对于分类处理,您应该使用 one-hot 编码。 做类似的事情

      formula <- dep_var ~ indep_var
      exp_mat <- model.matrix(formula, explanatory_matrix)
      solve(t(exp_mat) %*% exp_mat) %*% (t(exp_mat) %*% response_vec)
    
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-03-07
      • 1970-01-01
      • 2019-02-25
      • 2021-02-24
      • 2021-05-25
      • 2015-04-06
      • 1970-01-01
      • 2023-04-03
      相关资源
      最近更新 更多