【问题标题】:R: Application of group_by in dplyrR: group_by 在 dplyr 中的应用
【发布时间】:2016-04-10 17:58:43
【问题描述】:

我刚开始使用 dplyr,遇到以下两个问题,用group_by 应该很容易解决,但我不明白。 我的数据如下所示:

data <- data.frame(cbind("year" = c(2010, 2010, 2010, 2011, 2012, 2012, 2012, 2012),
                     "institution" = c("a", "a", "b", "a", "a", "a", "b", "b"),
                     "branch.num" = c(1, 2, 1, 1, 1, 2, 1, 2)))

data
#  year institution branch.num
#1 2010           a          1
#2 2010           a          2
#3 2010           b          1
#4 2011           a          1
#5 2012           a          1
#6 2012           a          2
#7 2012           b          1
#8 2012           b          2

数据是层次结构的:一个机构最高层可以有多个分支机构,从1开始编号。

问题 1:我想选择只包含分支机构的行,每年都有一个值,即示例数据中只有机构 a 的分支机构 1,所以选择应该是第 1、4 和 5 行。

Pronlem 2:我想知道一个机构多年来的平均分支机构数量。这就是机构 a (2+1+2)/3 = 1.67 和机构 b (1+0+2)/3 = 1 的示例。

【问题讨论】:

    标签: r group-by dplyr


    【解决方案1】:

    这是一种解决方案:

    问题 #1:

    library(dplyr)
    nYears <- n_distinct(data$year)
    data %>% group_by(institution, branch.num) %>% filter(n_distinct(year) == nYears)
    Source: local data frame [3 x 3]
    Groups: institution, branch.num [1]
    
        year institution branch.num
      (fctr)      (fctr)     (fctr)
    1   2010           a          1
    2   2011           a          1
    3   2012           a          1
    

    问题 #2:

    data %>% group_by(institution, year) %>% summarise(nBranches = n_distinct(branch.num)) %>% ungroup() %>% group_by(institution) %>% summarise(meanBranches = sum(nBranches)/nYears)
    Source: local data frame [2 x 2]
    
      institution meanBranches
           (fctr)        (dbl)
    1           a     1.666667
    2           b     1.000000
    

    【讨论】:

      猜你喜欢
      • 2021-04-17
      • 1970-01-01
      • 2019-12-06
      • 2017-06-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多