【发布时间】:2022-01-01 22:13:43
【问题描述】:
我有一个关于在包含交互作用(连续 * 因子)的 GLM 中变量的相对重要性的问题。
我正在尝试一种基于对解释的变体进行分区的方法,通过(伪)-R-squared 近似。但我不确定如何 (1) 在 GLM 中,以及 (2) 使用包含交互的模型。
为简单起见,我准备了一个带有单次交互的高斯 GLM 示例模型(使用 mtcars 数据集,请参阅帖子末尾的代码)。但我实际上有兴趣将该方法应用于可能包含多个交互的广义泊松 GLM。 测试模型提出了几个问题:
- 如何正确划分 R 平方?我尝试过划分,但不确定这是否正确。
- 每个项的 r-squared 与完整模型的 r-squared 不相加(甚至不接近)。这也发生在不包含交互作用的模型中。除了划分r平方的错误(我仍然认为自己是统计的新手:P);这也会受到共线性的影响吗?缩放连续预测变量后,方差膨胀因子低于 3(未缩放的模型具有最高的 VIF = 5.7)。
非常感谢任何帮助!
library(tidyverse)
library(rsq)
library(car)
data <- mtcars %>%
# scale reduces collinearity: without standardizing, the variance inflation factor for the factor is 5.7
mutate(disp = scale(disp))
data$am <- factor(data$am)
summary(data)
# test model, continuous response (miles per gallon), type of transmission (automatic/manual) as factor, displacement as continuous
model <-
glm(mpg ~ am + disp + am:disp,
data = data,
family = gaussian(link = "identity"))
drop1(model, test = "F")
# graph the data
ggplot(data = data, aes(x = disp, y = mpg, col = am)) + geom_jitter() + geom_smooth(method = "glm")
# Attempted partitioning
(rsq_full <- rsq::rsq(model, adj = TRUE, type = "v"))
(rsq_int <- rsq_full - rsq::rsq(update(model, . ~ . - am:disp), adj = TRUE, type = "v"))
(rsq_factor <- rsq_full - rsq::rsq(update(model, . ~ . - am - am:disp), adj = TRUE, type = "v"))
(rsq_cont <- rsq_full - rsq::rsq(update(model, . ~ . - disp - am:disp), adj = TRUE, type = "v"))
c(rsq_full, rsq_int + rsq_factor + rsq_cont)
car::vif(model)
# A simpler model with no interaction
model2 <- glm(mpg ~ am + disp, data = data, family = gaussian(link = "identity"))
drop1(model2, test = "F")
(rsq_full2 <- rsq::rsq(model2, adj = TRUE, type = "v"))
(rsq_factor2 <- rsq_full2 - rsq::rsq(update(model2, . ~ . - am), adj = TRUE, type = "v"))
(rsq_cont2 <- rsq_full2 - rsq::rsq(update(model2, . ~ . - disp), adj = TRUE,type = "v"))
c(rsq_full2, rsq_factor2 + rsq_cont2)
car::vif(model2)
【问题讨论】:
标签: r glm interaction variance