【发布时间】:2021-05-19 21:16:57
【问题描述】:
如果新列不存在,我一直在努力尝试添加它。我在这里找到了答案:Adding column if it does not exist。
但是,在我的问题中,我必须在 purrr 环境中使用它。我试图调整上述答案,但它不符合我的需求。
这是我正在处理的示例:
假设我有一个包含两个data.frames 的列表:
library(tibble)
A = tibble(
x = 1:5, y = 1, z = 2
)
B = tibble(
x = 5:1, y = 3, z = 3, w = 7
)
dt_list = list(A, B)
我要添加的栏目是w:
cols = c(w = NA_real_)
另外,如果我想添加不存在的列,我可以执行以下操作:
既然存在,就不添加列了:
B %>% tibble::add_column(!!!cols[!names(cols) %in% names(.)])
# A tibble: 5 x 4
x y z w
<int> <dbl> <dbl> <dbl>
1 5 3 3 7
2 4 3 3 7
3 3 3 3 7
4 2 3 3 7
5 1 3 3 7
在这种情况下,由于它不存在,所以添加w:
A %>% tibble::add_column(!!!cols[!names(cols) %in% names(.)])
# A tibble: 5 x 4
x y z w
<int> <dbl> <dbl> <dbl>
1 1 1 2 NA
2 2 1 2 NA
3 3 1 2 NA
4 4 1 2 NA
5 5 1 2 NA
我尝试使用purrr 复制它(我不想使用 for 循环):
dt_list_2 = dt_list %>%
purrr::map(
~dplyr::select(., -starts_with("x")) %>%
~tibble::add_column(!!!cols[!names(cols) %in% names(.)])
)
但是输出和单独做是不一样的。
注意:这是我真正的问题的一个例子。事实上,我正在使用purrr 读取许多 *.csv 文件,然后应用一些数据转换。像这样的:
re_file <- list.files(path = dir_path, pattern = "*.csv")
cols_add = c(UCI = NA_real_)
file_list = re_file %>%
purrr::map(function(file_name){ # iterate through each file name
read_csv(file = paste0(dir_path, "//",file_name), skip = 2)
}) %>%
purrr::map(
~dplyr::select(., -starts_with("Textbox")) %>%
~dplyr::tibble(!!!cols[!names(cols) %in% names(.)])
)
【问题讨论】:
-
purrr::map(dt_list, function(x) {x$w <- x$w %||% NA; x})其中%||%来自那些整洁的包之一