【发布时间】:2016-09-25 20:57:57
【问题描述】:
我有一个训练集,其中 x 列代表正在进行比赛的特定体育场。显然,这些列在训练集中是线性相关的,因为一场比赛必须发生在至少一个体育场内。
但是我遇到的问题是,如果我通过了测试数据,它可能包括训练数据中没有看到的体育场。因此,我想在训练 R glm 时包括所有 x 列,以使每个体育场系数的平均值为零。然后,如果看到一个新的体育场,它基本上会得到所有体育场系数的平均值。
我遇到的问题是 R glm 函数似乎检测到我的训练集中有线性相关的列,并将其中一个系数设置为 NA 以使其余的列线性独立。我该怎么做:
停止 R 在 glm 函数中插入我的一列的 NA 系数并确保所有体育场系数总和为 0?
一些示例代码
# Past observations
outcome = c(1 ,0 ,0 ,1 ,0 ,1 ,0 ,0 ,1 ,0 ,1 )
skill = c(0.1,0.5,0.6,0.3,0.1,0.3,0.9,0.6,0.5,0.1,0.4)
stadium_1 = c(1 ,1 ,0 ,0 ,0 ,0 ,0 ,0 ,0 ,0 ,0 )
stadium_2 = c(0 ,0 ,1 ,1 ,1 ,1 ,1 ,0 ,0 ,0 ,0 )
stadium_3 = c(0 ,0 ,0 ,0 ,0 ,0 ,0 ,1 ,1 ,1 ,1 )
train_glm_data = data.frame(outcome, skill, stadium_1, stadium_2, stadium_3)
LR = glm(outcome ~ . - outcome, data = train_glm_data, family=binomial(link='logit'))
print(predict(LR, type = 'response'))
# New observations (for a new stadium we have not seen before)
skill = c(0.1)
stadium_1 = c(0 )
stadium_2 = c(0 )
stadium_3 = c(0 )
test_glm_data = data.frame(outcome, skill, stadium_1, stadium_2, stadium_3)
print(predict(LR, test_glm_data, type = 'response'))
# Note that in this case, the observation is simply the same as if we had observed stadium_3
# Instead I would like it to be an average of all the known stadiums coefficients
# If they all sum to 0 this is essentially already done for me
# However if not then the stadium_3 coefficient is buried somewhere in the intercept term
【问题讨论】:
-
您要适配的
glm型号是什么?如果您提供一个最小的可重现示例,将会有所帮助。