【问题标题】:Relative importance/Variation partitioning in a GLM containing an interaction包含交互的 GLM 中的相对重要性/变化分区
【发布时间】:2022-01-01 22:13:43
【问题描述】:

我有一个关于在包含交互作用(连续 * 因子)的 GLM 中变量的相对重要性的问题。

我正在尝试一种基于对解释的变体进行分区的方法,通过(伪)-R-squared 近似。但我不确定如何 (1) 在 GLM 中,以及 (2) 使用包含交互的模型。

为简单起见,我准备了一个带有单次交互的高斯 GLM 示例模型(使用 mtcars 数据集,请参阅帖子末尾的代码)。但我实际上有兴趣将该方法应用于可能包含多个交互的广义泊松 GLM。 测试模型提出了几个问题:

  1. 如何正确划分 R 平方?我尝试过划分,但不确定这是否正确。
  2. 每个项的 r-squared 与完整模型的 r-squared 不相加(甚至不接近)。这也发生在不包含交互作用的模型中。除了划分r平方的错误(我仍然认为自己是统计的新手:P);这也会受到共线性的影响吗?缩放连续预测变量后,方差膨胀因子低于 3(未缩放的模型具有最高的 VIF = 5.7)。

非常感谢任何帮助!


library(tidyverse)
library(rsq)
library(car)

data <- mtcars %>%
  # scale reduces collinearity: without standardizing, the variance inflation factor for the factor is 5.7
  mutate(disp = scale(disp))
data$am <- factor(data$am)

summary(data)

# test model, continuous response (miles per gallon), type of transmission (automatic/manual) as factor, displacement as continuous
model <-
  glm(mpg ~ am + disp + am:disp,
      data = data,
      family = gaussian(link = "identity"))
drop1(model, test = "F")

# graph the data
ggplot(data = data, aes(x = disp, y = mpg, col = am)) + geom_jitter() + geom_smooth(method = "glm")

# Attempted partitioning
(rsq_full <- rsq::rsq(model, adj = TRUE, type = "v"))

(rsq_int <- rsq_full - rsq::rsq(update(model, . ~ . - am:disp), adj = TRUE, type = "v"))

(rsq_factor <- rsq_full - rsq::rsq(update(model, . ~ . - am - am:disp), adj = TRUE, type = "v"))

(rsq_cont <- rsq_full - rsq::rsq(update(model, . ~ . - disp - am:disp), adj = TRUE, type = "v"))

c(rsq_full, rsq_int + rsq_factor + rsq_cont)

car::vif(model)


# A simpler model with no interaction
model2 <- glm(mpg ~ am + disp, data = data, family = gaussian(link = "identity"))
drop1(model2, test = "F")

(rsq_full2 <- rsq::rsq(model2, adj = TRUE, type = "v"))
(rsq_factor2 <- rsq_full2 - rsq::rsq(update(model2, . ~ . - am), adj = TRUE, type = "v"))
(rsq_cont2 <- rsq_full2 - rsq::rsq(update(model2, . ~ . - disp), adj = TRUE,type = "v"))

c(rsq_full2, rsq_factor2 + rsq_cont2)

car::vif(model2)


【问题讨论】:

    标签: r glm interaction variance


    【解决方案1】:

    给定:

    1. y = A + B + A * B

    我会将它的 R 平方值与其简单版本的值进行比较:

    1. y = A + B
    2. y = A
    3. y = B

    如果没有交互,我希望

    r-squared(model1) = r-squared(model2)
    

    这应该适用于任何线性模型。即使存在交互,它也应该有助于比较预测变量的主要影响。我知道这是有争议的,但是如果您查看下图所示的场景,则仅在考虑到预测器 B 时,预测器 A 才具有信息性;相反,预测器 B 本身也具有一定的预测能力(B1 的 y 高于 B2 的 y,无论它们属于哪个级别的 A)。

    这是一个模拟数据示例(避免共线性和非正态性问题):

    # simulate data:
    df <- data.frame(Species = as.factor(c(rep("Species A", 200),
                                rep("Species B", 200)
                                )),
                    Treatment = as.factor(rep(c("diet 1", "diet 2","diet 1", "diet 2"), each=100)),
                    body.weight = c(rnorm(n=100, 30, 5),
                                    rnorm(n=100, 29.9, 5),
                                    rnorm(n=100, 55, 5),
                                    rnorm(n=100, 90, 5)
                                    )
                    )
    

    # Let's fit and compare the alternative models:
    lm.interactive <- lm(body.weight ~ Species * Treatment, data=df)
    lm.additive <- lm(body.weight ~ Species + Treatment, data=df)
    lm.only.species <- lm(body.weight ~ Species, data=df)
    lm.only.Treatment <- lm(body.weight ~ Treatment, data=df)
    lm.null <- lm(body.weight ~ 1, data=df)
    
    # obtain R^2:
    
    
    summary(lm.only.Treatment)$adj.r.squared # main effect of Treatment
    summary(lm.only.species)$adj.r.squared # main effect of species ID. 
    # As the figure suggests, it's larger than the main effect of Treatment 
    # (species identity affects body weight regardless of treatment)
    summary(lm.additive)$adj.r.squared # sum of the main effects
    summary(lm.interactive)$adj.r.squared # main effects + interaction
    
    # fraction of variance explained by the interaction alone:
    summary(lm.interactive)$adj.r.squared - summary(lm.additive)$adj.r.squared
    

    我不确定我们是否真的可以谈论“由交互单独解释的方差分数”。谈论由于包含在交互项中而增加的解释方差可能更合适。

    我不确定我建议的方法在统计上的合理性如何,它的局限性,或者它对于不平衡的数据集是否可靠地工作。这种方法的一个问题是,R 平方的差异无法进行统计测试,因为我们每个模型只有一个 R 平方值。解决它的一种方法是使用自举获得每个模型的 R 平方值分布。

    【讨论】:

    • 谢谢,这看起来是一种明智的做法。我想知道在多元方法(假设超过 3 个解释变量)中,当从模型中删除一个术语(而不是单变量模型)时,计算 R^2 的下降是否更有意义。顺便说一句,在您的示例中,R^2 加起来非常好,但在我的示例中,R^2 不加起来......我认为它与 GLM 方法(不是 OLS)有关,并且存在共线性(?)。另外,为迟到的回复道歉!干杯
    • 如果您的两个或多个预测变量是连续的,则绝对有可能存在共线性。不过我不知道怎么解决。
    • ...我不知道如何解决共线性问题,除非顺序删除涉及共线性的预测变量。请参阅“第 5 步:协变量之间是否存在共线性?”在 doi 中:10.1111/j.2041-210X.2009.00001.x 和参考中的更多细节。如果两个预测变量共线,则对其主要影响的估计必然会受到它的影响。我怀疑这同样适用于共线预测变量之间的任何交互。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2023-03-03
    • 2014-09-17
    • 2020-07-12
    • 2020-05-05
    • 1970-01-01
    • 1970-01-01
    • 2010-09-28
    相关资源
    最近更新 更多