【问题标题】:Conditional filtering based on the level of a factor R基于因子 R 水平的条件过滤
【发布时间】:2014-09-01 00:42:33
【问题描述】:

我想清理以下代码。具体来说,我想知道是否可以合并三个过滤器语句,以便最终得到包含数据行“spring”(如果存在)的最终 data.frame(rind()),数据行“如果“spring”不存在,则为fall”,如果“spring”和“fall”都不存在,则最后是数据行。下面的代码看起来非常笨重且效率低下。我试图让自己摆脱 for(),所以希望解决方案不会涉及到一个。这可以使用 dplyr 完成吗?

# define a %not% to be the opposite of %in%
library(dplyr)
`%not%` <- Negate(`%in%`)
f <- c("a","a","a","b","b","c")
s <- c("fall","spring","other", "fall", "other", "other")
v <- c(3,5,1,4,5,2)
(dat0 <- data.frame(f, s, v))
sp.tmp <- filter(dat0, s == "spring")
fl.tmp <- filter(dat0, f %not% sp.tmp$f, s == "fall")
ot.tmp <- filter(dat0, f %not% sp.tmp$f, f %not% fl.tmp$f, s == "other")
rbind(sp.tmp,fl.tmp,ot.tmp)

【问题讨论】:

  • 在“a”、“b”或“c”中是否有可能存在多个“弹簧”?如果是,你想保留所有这些还是只保留第一个?
  • “a”、“b”、...不能有多个“弹簧”

标签: r dplyr


【解决方案1】:

看起来在f 的每组中,您想要提取spring、fall 或other 的行,按优先级降序排列。

如果您首先将您的偏好排序作为实际因素排序:

dat0$s <- factor(dat0$s, levels=c("spring", "fall", "other"))

然后您可以使用this dplyr solution 获取每个组中的最小行(相对于该因子):

newdat <- dat0 %.% group_by(f) %.% filter(rank(s) == 1)

【讨论】:

  • 谢谢,这正是我想要的。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-12-28
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多