【问题标题】:How to collapse rows with same identifier and retain non-empty column values?如何折叠具有相同标识符的行并保留非空列值?
【发布时间】:2019-05-07 01:11:30
【问题描述】:

我有一个表(经过一些初始处理后)有多行具有相同的主标识符但具有不同的列值(0 或值 > 0)。

示例表 主标识符为“produce”

df = data.frame(produce = c("apples","apples", "bananas","bananas"),
                grocery1=c(0,1,1,1),
                grocery2=c(1,0,1,1),
                grocery3=c(0,0,1,1))


###########################

> df
  produce grocery1 grocery2 grocery3
1  apples        0        1        0
2  apples        1        0        0
3 bananas        1        1        1
4 bananas        1        1        1

我想折叠(或合并?)具有相同标识符的行,并在每列中保留非空(此处为任何非零值)值

所需输出示例

 shopping grocery1 grocery2 grocery3
1   apples        1        1        0
2  bananas        1        1        1

tidyverse 中是否有我缺少的简单函数或管道可以处理这个问题?

【问题讨论】:

  • dplyr::group_by() 加上dplyr::summarise()

标签: r dplyr tidyr


【解决方案1】:

使用基础 R aggregate 我们可以做到

aggregate(.~produce, df, function(x) +any(x > 0))

#  produce grocery1 grocery2 grocery3
#1  apples        1        1        0
#2 bananas        1        1        1

或使用dplyr

library(dplyr)
df %>%
  group_by(produce) %>%
  summarise_all(~+any(. > 0))

#  produce grocery1 grocery2 grocery3
#  <fct>      <int>    <int>    <int>
#1 apples         1        1        0
#2 bananas        1        1        1

和data.table一样

library(data.table)
setDT(df)[, lapply(.SD, function(x) +any(x > 0)), by=produce]

【讨论】:

  • 谢谢!操作表格对我来说还不直观(无论是在 tidyverse 内部还是外部)......希望在这里和我自己的更多循环会改变这一点。
  • 这不只是将max 应用为聚合函数吗?
  • hmmm..不这么认为。也许你是对的,重新阅读我不清楚 OP 的目标以及如果有任何值 > 1 他们会期望输出的问题。
【解决方案2】:

我们可以使用max

library(dplyr)
df %>%
   group_by(produce) %>% 
   summarise_all(max)
# A tibble: 2 x 4
#  produce grocery1 grocery2 grocery3
#  <fct>      <dbl>    <dbl>    <dbl>
#1 apples         1        1        0
#2 bananas        1        1        1

【讨论】:

    猜你喜欢
    • 2020-01-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-07-11
    • 1970-01-01
    • 2019-04-02
    • 1970-01-01
    相关资源
    最近更新 更多