【问题标题】:Include linearly dependent features in glm在 glm 中包含线性相关的特征
【发布时间】:2016-09-25 20:57:57
【问题描述】:

我有一个训练集,其中 x 列代表正在进行比赛的特定体育场。显然,这些列在训练集中是线性相关的,因为一场比赛必须发生在至少一个体育场内。

但是我遇到的问题是,如果我通过了测试数据,它可能包括训练数据中没有看到的体育场。因此,我想在训练 R glm 时包括所有 x 列,以使每个体育场系数的平均值为零。然后,如果看到一个新的体育场,它基本上会得到所有体育场系数的平均值。

我遇到的问题是 R glm 函数似乎检测到我的训练集中有线性相关的列,并将其中一个系数设置为 NA 以使其余的列线性独立。我该怎么做:

停止 R 在 glm 函数中插入我的一列的 NA 系数并确保所有体育场系数总和为 0?

一些示例代码

# Past observations
outcome   = c(1  ,0  ,0  ,1  ,0  ,1  ,0  ,0  ,1  ,0  ,1  )
skill     = c(0.1,0.5,0.6,0.3,0.1,0.3,0.9,0.6,0.5,0.1,0.4)
stadium_1 = c(1  ,1  ,0  ,0  ,0  ,0  ,0  ,0  ,0  ,0  ,0  )
stadium_2 = c(0  ,0  ,1  ,1  ,1  ,1  ,1  ,0  ,0  ,0  ,0  )
stadium_3 = c(0  ,0  ,0  ,0  ,0  ,0  ,0  ,1  ,1  ,1  ,1  )

train_glm_data = data.frame(outcome, skill, stadium_1, stadium_2,     stadium_3)
LR = glm(outcome ~ . - outcome, data = train_glm_data,  family=binomial(link='logit'))
print(predict(LR, type = 'response'))

# New observations (for a new stadium we have not seen before)
skill     = c(0.1)
stadium_1 = c(0  )
stadium_2 = c(0  )
stadium_3 = c(0  )

test_glm_data = data.frame(outcome, skill, stadium_1, stadium_2, stadium_3)
print(predict(LR, test_glm_data, type = 'response'))

# Note that in this case, the observation is simply the same as if we had observed stadium_3
# Instead I would like it to be an average of all the known stadiums coefficients
# If they all sum to 0 this is essentially already done for me
# However if not then the stadium_3 coefficient is buried somewhere in the intercept term

【问题讨论】:

标签: r glm


【解决方案1】:

要估计所有虚拟变量的系数,您可以在公式中添加“-1”,这将删除截距:

LR = glm(outcome ~ . - outcome - 1, data = train_glm_data, family=binomial(link='logit'))

系数:

coef(LR)
#      skill  stadium_1  stadium_2  stadium_3 
# -2.8080177  0.8424053  0.7541226  1.1313135 

对于看不见的训练水平问题,@hack-r 提出了一些好主意。另一个想法是将1/n(其中n 是观察到的体育场的数量)估算为新观察的所有虚拟变量。

【讨论】:

    【解决方案2】:
    train_glm_data$stadium <- NA
    train_glm_data$stadium[train_glm_data$stadium_1==1] <- "Stadium 1"
    train_glm_data$stadium[train_glm_data$stadium_2==1] <- "Stadium 2"
    train_glm_data$stadium[train_glm_data$stadium_3==1] <- "Stadium 3"
    train_glm_data$stadium_1 <- NULL
    train_glm_data$stadium_2 <- NULL
    train_glm_data$stadium_3 <- NULL
    
    train_glm_data$stadium         <- as.factor(train_glm_data$stadium)
    levels(train_glm_data$stadium) <- c("Stadium 1", "Stadium 2", "Stadium 3", "Stadium 4")
    train_glm_data                 <- rbind(train_glm_data, c(
                                          round(mean(outcome)), mean(skill),
                                          "Stadium 4"
                                        ))
    train_glm_data$outcome <- as.numeric(train_glm_data$outcome)
    train_glm_data$skill   <- as.numeric(train_glm_data$skill)
    LR = glm(outcome ~ stadium + skill, data = train_glm_data,  family=binomial(link='logit'))
    print(predict(LR, type = 'response'))
    
    # New observations (for a new stadium we have not seen before)
    skill     = c(0.1)
    stadium   = "Stadium 4"
    
    test_glm_data = data.frame(skill, stadium)
    print(predict(LR, test_glm_data, type = 'response'))
    

    关于如何包含所有级别的系数的问题——不要这样做。它被称为虚拟变量陷阱。如果不排除参考水平,则数据矩阵变为奇异。

    唯一的例外是您估计 no-intercept 模型。阅读更多关于虚拟变量陷阱here.

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2022-07-05
      • 2021-12-24
      • 2022-07-06
      • 1970-01-01
      • 2022-01-01
      • 2019-04-25
      • 2020-09-27
      相关资源
      最近更新 更多