【问题标题】:R: Create new column based list of values from a multiple columnsR:从多列创建新的基于列的值列表
【发布时间】:2020-02-09 01:48:29
【问题描述】:

我想根据多个列中存在的列表中的任何值创建一个新列 (T/F)。对于这个示例,我使用 mtcars 作为示例,在两列中搜索两个值,但我的实际挑战是多列中的多个值。

我有一个使用下面包含的filter_at() 的成功过滤器,但我无法将该逻辑应用于变异:

# there are 7 cars with 6 cyl
mtcars %>%
  filter(cyl == 6)

# there are 2 cars with 19.2 mpg, one with 6 cyl, one with 8
mtcars %>% 
  filter(mpg == 19.2)

# there are 8 rows with either.
# these are the rows I want as TRUE
mtcars %>% 
  filter(mpg == 19.2 | cyl == 6)

# set the cols to look at
mtcars_cols <- mtcars %>% 
  select(matches('^(mp|cy)')) %>% names()

# set the values to look at
mtcars_numbs <- c(19.2, 6)

# result is 8 vars with either value in either col.
# this is a successful filter of the data
out1 <- mtcars %>% 
    filter_at(vars(mtcars_cols), any_vars(
        . %in% mtcars_numbs
        )
      )

# shows set with all 6 cyl, plus one 8cyl 21.9 mpg
out1 %>% 
  select(mpg, cyl)

# This attempts to apply the filter list to the cols,
# but I only get 6 rows as True
# I tried to change == to %in& but that results in an error
out2 <- mtcars %>%
    mutate(
      myset = rowSums(select(., mtcars_cols) == mtcars_numbs) > 0
    )

# only 6 rows returned
out2 %>% 
  filter(myset == T)

我不确定为什么会跳过这两行。我认为可能是rowSums 的使用以某种方式聚合了这两行。

【问题讨论】:

    标签: r dplyr


    【解决方案1】:

    如果我们想做相应的检查,使用map2可能会更好

     library(dplyr)
     library(purrr)
     map2_df(mtcars_cols, mtcars_numbs, ~ 
           mtcars %>%
               filter(!! rlang::sym(.x) == .y)) %>%
         distinct
    

    注意:与浮点数进行比较 (==) 可能会遇到麻烦,因为精度可能会有所不同并导致 FALSE


    另外,请注意== 仅在lhs 和rhs 元素具有相同长度或rhs 向量为length 1 时才有效(此处发生回收)。如果length 大于 1 且不等于 lhs 向量的长度,则回收将按列顺序进行比较。

    我们可以replicate 使长度相等,现在它应该可以工作了

    mtcars %>%
     mutate(
       myset = rowSums(select(., mtcars_cols) == mtcars_numbs[col(select(., mtcars_cols))]) > 0
       ) %>% pull(myset) %>% sum
    #[1] 8
    

    在上面的代码中,select 使用了两次以便更好地理解。否则,我们也可以使用rep

    mtcars %>%
     mutate(
       myset = rowSums(select(., mtcars_cols) == rep(mtcars_numbs, each = n())) > 0
        ) %>% 
       pull(myset) %>%
       sum
    #[1] 8
    

    【讨论】:

    • 关于浮点数的好点,虽然我在我的真实数据集中搜索字符串而不是数字......我应该使用不同的示例数据集。今天我会试一试,让你知道。我的另一个想法是构建一个函数来循环遍历值的列列表并使用 mutate 设置答案。
    • 好的,我承认我在尝试将map2 概念应用于我的数据时有点迷茫。我也觉得使用 mtcars 并且浮点数据与我的实际挑战并不真正相似,所以我对我所追求的和我尝试过的东西有一个很长的解释at this published notebook。我将在上面添加一些修改以反映这一点。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-12-14
    • 1970-01-01
    • 1970-01-01
    • 2022-08-11
    • 1970-01-01
    相关资源
    最近更新 更多