【问题标题】:Filtering all columns depending on one column using dplyr使用 dplyr 根据一列过滤所有列
【发布时间】:2017-04-15 09:30:34
【问题描述】:

我想使用 dplyr 过滤至少一列(不包括 P)大于 P 的行。试图找出一个过滤所有列的解决方案。

例子

library(dplyr)

df <-  tibble(P = c(2,4,5,6,1.4), B = 
                c(2.1,3,5.5,1.2, 2), 
                C = c(2.2, 3.8, 5.7, 5,
                  1.5))

期望的输出

df <- filter(df, B > P | C > P)
df

一种使用 apply 的解决方案,如果可能,我想避免:

filter(df, apply(df, 1, function(x) sum(x > x[1]) > 1))

【问题讨论】:

  • 为什么需要data.table?
  • 你是对的,因为 tibble 来自 dplyr,所以不需要。
  • 你可以简单地做df[rowSums(df[-1] &gt; df$P)&gt;0,]
  • 我相信@Sotos 的矢量化解决方案会很快。 filter 基于使用数据列,而不是数据行。这就是为什么你很难把它融入dplyr。
  • 你也可以和slice一起使用as shown here

标签: r dplyr


【解决方案1】:

没有dplyr...

df2 <- df[df$P!=apply(df,1,max),]

或dplyr...

df3 <- df %>% filter(P!=apply(df,1,max))

【讨论】:

  • 谢谢,我试图避免使用 apply 的解决方案,将编辑我的问题。
【解决方案2】:

这是一个使用tidyverse 的选项,我们利用purrr 中的map 和reduce 函数来获得逻辑vector 到extract(来自magrittr)的行原始数据集

library(tidyverse)
library(magrittr)
df %>% 
    select(-one_of("P")) %>% 
    map(~ .> df$P) %>% 
    reduce(`|`) %>%
    extract(df, .,)
# A tibble: 3 × 3
#      P     B     C
#  <dbl> <dbl> <dbl>
#1   2.0   2.1   2.2
#2   5.0   5.5   5.7
#3   1.4   2.0   1.5

这也可以使用dplyr(即将发布0.6.0)的开发版本转换为函数,其中引入了quosures和unquote进行评估。 enquo 与base R 中的substitute 几乎相似,后者接受用户输入并将其转换为quosure,one_of 接受字符串参数,因此可以使用quo_name 将其转换为字符串

funFilter <- function(dat, colToCompare){
   colToCompare <- quo_name(enquo(colToCompare))

   dat %>%
       select(-one_of(colToCompare)) %>%
       map(~ .> dat[[colToCompare]]) %>%
       reduce(`|`) %>%
       extract(dat, ., )
}

funFilter(df, P)#compare all other columns with P
# A tibble: 3 × 3
#      P     B     C
#  <dbl> <dbl> <dbl>
#1   2.0   2.1   2.2
#2   5.0   5.5   5.7
#3   1.4   2.0   1.5

funFilter(df, B) #compare all other columns with B
# A tibble: 4 × 3
#      P     B     C
#  <dbl> <dbl> <dbl>
#1     2   2.1   2.2
#2     4   3.0   3.8
#3     5   5.5   5.7
#4     6   1.2   5.0

我们也可以解析表达式

v1 <- setdiff(names(df), "P")
filter(df, !!rlang::parse_quosure(paste(v1, "P", sep=" > ", collapse=" | ")))
# A tibble: 3 × 3
#     P     B     C
#    <dbl> <dbl> <dbl>
#1   2.0   2.1   2.2
#2   5.0   5.5   5.7
#3   1.4   2.0   1.5

这也可以做成函数

funFilter2 <- function(dat, colToCompare){
    colToCompare <- quo_name(enquo(colToCompare))
    v1 <- setdiff(names(dat), colToCompare)
    expr <- rlang::parse_quosure(paste(v1, colToCompare, sep= " > ", collapse= " | "))
    dat %>%
        filter(!!expr)
}

funFilter2(df, P)
# A tibble: 3 × 3
#      P     B     C
#  <dbl> <dbl> <dbl>
#1   2.0   2.1   2.2
#2   5.0   5.5   5.7
#3   1.4   2.0   1.5

funFilter2(df, B)
# A tibble: 4 × 3
#      P     B     C
#  <dbl> <dbl> <dbl>
#1     2   2.1   2.2
#2     4   3.0   3.8
#3     5   5.5   5.7
#4     6   1.2   5.0

或者另一种方法是pmax

df %>%
   filter(do.call(pmax, .) > P)
# A tibble: 3 × 3
#      P     B     C
#   <dbl> <dbl> <dbl>
#1   2.0   2.1   2.2
#2   5.0   5.5   5.7
#3   1.4   2.0   1.5

【讨论】:

  • 毫无疑问 dplyr 不错,但有些变种有受虐倾向。
  • @DieterMenne 我认为您对获得输出的步骤数是正确的。基于apply 或rowSums 的方法会更紧凑,但它仍然不是一个tidyverse 语法。
猜你喜欢
  • 2021-11-11
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-03-05
  • 2020-12-08
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多