【问题标题】:Issues with cumsum in RR中的cumsum问题
【发布时间】:2021-04-07 18:59:55
【问题描述】:

这里是示例数据和包。我正在使用的代码如下。它适用于前四行,但之后出现问题。想要的结果在最底部。我需要 cumsum 只查看区域、周期组合...... 001 和 2020q1。在这种情况下,将有 4 个分组(001/2020q1、003/2020q1、001/2020q2、003/2020q2)。我将如何进行这样的过程?我有一种感觉,我在 group by 子句中遗漏了一些东西,但到目前为止还在绕圈子。

这是上一个问题的延续。这有更多的数据,并且涉及更多。

 library(readxl)
 library(dplyr)
 library(data.table)
 library(odbc)
 library(DBI)
 library(stringr)

employment <- c(1,45,125,130,165,260,2,46,127,132,167,265,50,61,110,121,170,305,55,66,112,123,172,310)
small <- c(1,1,2,2,3,4,1,1,2,2,3,4,1,1,2,2,3,4,1,1,2,2,3,4)
area <-c(001,001,001,001,001,001,001,001,001,001,001,001,003,003,003,003,003,003,003,003,003,003,003,003)
year<-c(2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020,2020)
qtr <-c(1,1,1,1,1,1,2,2,2,2,2,2,1,1,1,1,1,1,2,2,2,2,2,2)

smbtest <- data.frame(employment,small,area,year,qtr)


 smbsummary2<-smbtest %>% 
 mutate(period = paste0(year,"q",qtr)) %>%
 select(area,period,employment,small) %>%
 group_by(area,period,small) %>%
 summarise(employment = sum(employment), worksites = n(), 
        .groups = 'drop') %>% 
 mutate(employment = cumsum(employment),
     worksites = cumsum(worksites))


area    period     small    employment    worksites
 001     2020q1     1          46            2
 001     2020q1     2          303           4
 001     2020q1     3          466           5
 001     2020q1     4          726           6
 003     2020q1     1          48            2
 003     2020q1     2          307           4
 003     2020q1     3          474           5
 003     2020q1     4          739           6
 001     2020q2     1          111           2
 001     2020q2     2          342           4
 001     2020q2     3          512           5
 001     2020q1     4          817           6
 and so on. 

【问题讨论】:

  • 你为什么使用.groups='drop'?听起来您想保留区域/期间。在总结 或 后添加另一个 group_by() 语句或更改 .groups 选项? (评论而不是回答,因为我在猜测 - 没有仔细查看代码)
  • @BenBolker,这是在之前的问题中提出的。我对此还是有点陌生​​。

标签: r dplyr group-by cumsum


【解决方案1】:

.groups = 'drop' 删除所有组,而我们需要.groups = 'drop_last'。根据显示的预期输出,应该删除“小”列。默认情况下,summarise 执行.groups = 'drop_last,如果我们想指定它来删除警告,可以这样做

smbsummary2 <- smbtest %>% 
 mutate(period = paste0(year,"q",qtr)) %>%
 select(area,period,employment,small) %>%
 group_by(area,period,small) %>%
 summarise(employment = sum(employment), worksites = n(), 
        .groups = 'drop_last') %>%  mutate(employment = cumsum(employment),
     worksites = cumsum(worksites))

【讨论】:

    猜你喜欢
    • 2021-07-27
    • 2022-01-24
    • 1970-01-01
    • 1970-01-01
    • 2015-02-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多