【发布时间】:2022-01-11 21:04:28
【问题描述】:
我想用 R 中的控制函数(使用 @987654323 @ 和 broom)。我想基于具有因变量y、内生变量x、此内生变量z1 的工具和外生变量z2 的分组数据框来实现这一点。遵循两阶段最小二乘法 (2SLS) 方法,我将运行:(1) 在 z1 和 z2 上回归 x 和 (2) 在 x 上回归 y, z2 和 v(来自 (1) 的残差)。有关此方法的更多详细信息,请参阅:https://www.irp.wisc.edu/newsevents/workshops/appliedmicroeconometrics/participants/slides/Slides_14.pdf。不幸的是,我无法在没有错误的情况下运行第二次回归(见下文)。
我的数据如下所示:
df <- data.frame(
id = sort(rep(seq(1, 20, 1), 5)),
group = rep(seq(1, 4, 1), 25),
y = runif(100),
x = runif(100),
z1 = runif(100),
z2 = runif(100)
)
其中id 是观察的标识符,group 是组的标识符,其余部分在上面定义。
library(tidyverse)
library(broom)
# Nest the data frame
df_nested <- df %>%
group_by(group) %>%
nest()
# Run first stage regression and retrieve residuals
df_fit <- df_nested %>%
mutate(
fit1 = map(data, ~ lm(x ~ z1 + z2, data = .x)),
resids = map(fit1, residuals)
)
现在,我想运行第二阶段回归。我已经尝试了两件事。
第一:
df_fit %>%
group_by(group) %>%
unnest(c(data, resids)) %>%
do(lm(y ~ x + z2, data = .x))
这会产生Error in is.data.frame(data) : object '.x' not found。
第二:
df_fit %>%
mutate(
fit2 = map2(data, resids, ~ lm(y ~ x + z2, data = .x))
)
df_fit %>% unnest(fit2)
这会产生:Error: Must subset columns with a valid subscript vector. x Subscript has the wrong type `grouped_df< 。如果您要处理更大的数据集,则第二种方法甚至会遇到存储问题。
这是如何正确完成的?
【问题讨论】:
-
我以更一般的方式重新表述了上述问题(重点是在最终回归中包含来自先前回归的元素)。你可以在这里找到它:stackoverflow.com/questions/70287136/….
标签: r regression tidyverse broom