【问题标题】:R rbind and group by using dplyrR rbind 并使用 dplyr 进行分组
【发布时间】:2019-10-24 15:26:36
【问题描述】:

我有以下数据

library(dplyr)

df1 <- tibble(
year = c("2001","2001", "2001", "2001", "2002","2002", "2002", "2002"),
type = c("Animals", "Animals", "People", "People", "Animals", "Animals", "People", "People"),
type_group = c("Dogs", "Cats", "John", "Jane", "Dogs", "Cats", "John", "Jane"),
analysis1 = c(32.7, 67.5, 34.6, 56.5, 56.7, 78.5, 98.9, 87.3),
analysis2 = c(23.7, 89.4, 45.8, 98.6, 45.7, 45.7, 23.6, 23.6),
analysis3 = c(45.7, 45.7, 23.6, 23.6, 14.4, 45.4, 98.0, 12.2),
analysis4 = c(14.4, 45.4, 98.0, 12.2, 34.6, 44.3, 23.8, 16.3))

我正在使用rbind 创建新行,其中包含一些新的计算,您将在下面的代码中看到这些计算。

我想知道是否有一种更简洁、更快捷的方法来执行此操作。我确定一定有……我的数据有大约 30 年的时间和大约 60 个变量,所以要使用我在这里开发的示例需要很长时间才能在我的真实数据的脚本中编写:

df1 %>% 
  filter(year =="2001") %>% 
rbind(c("2001", "People diff","John and Jane", 
            df1$analysis1[df1$type_group == 'John'] - df1$analysis1[df1$type_group == 'Jane'],
            df1$analysis2[df1$type_group == 'John'] - df1$analysis2[df1$type_group == 'Jane'],
            df1$analysis3[df1$type_group == 'John'] - df1$analysis3[df1$type_group == 'Jane'],
            df1$analysis4[df1$type_group == 'John'] - df1$analysis4[df1$type_group == 'Jane'])) %>% 
  rbind(c("2001","Animals diff","Dogs and cats", 
            df1$analysis1[df1$type_group == 'Cats'] - df1$analysis1[df1$type_group == 'Dogs'],
            df1$analysis2[df1$type_group == 'Cats'] - df1$analysis2[df1$type_group == 'Dogs'],
            df1$analysis3[df1$type_group == 'Cats'] - df1$analysis3[df1$type_group == 'Dogs'],
            df1$analysis4[df1$type_group == 'Cats'] - df1$analysis4[df1$type_group == 'Dogs'])) -> data_2001


df1 %>% 
  filter(year =="2002") %>% 
  rbind(c("2002", "People diff","John and Jane", 
          df1$analysis1[df1$type_group == 'John'] - df1$analysis1[df1$type_group == 'Jane'],
          df1$analysis2[df1$type_group == 'John'] - df1$analysis2[df1$type_group == 'Jane'],
          df1$analysis3[df1$type_group == 'John'] - df1$analysis3[df1$type_group == 'Jane'],
          df1$analysis4[df1$type_group == 'John'] - df1$analysis4[df1$type_group == 'Jane'])) %>% 
  rbind(c("2002","Animals diff","Dogs and cats", 
          df1$analysis1[df1$type_group == 'Cats'] - df1$analysis1[df1$type_group == 'Dogs'],
          df1$analysis2[df1$type_group == 'Cats'] - df1$analysis2[df1$type_group == 'Dogs'],
          df1$analysis3[df1$type_group == 'Cats'] - df1$analysis3[df1$type_group == 'Dogs'],
          df1$analysis4[df1$type_group == 'Cats'] - df1$analysis4[df1$type_group == 'Dogs'])) -> data_2002

rbind(data_2001, data_2002) -> final_data

感谢任何帮助!谢谢

【问题讨论】:

    标签: r dplyr tidyr reshape2


    【解决方案1】:

    首先,我认为您的分析是不正确的,除非它是这样设计的。在您的rbind 中,您使用df1$analysis1[df1$type_group == 'John'] 包含两年的数据,但仅将它们绑定一年并调用例如2001.

    快速简便的方法是使用tidyr 包中的spread 和gather,例如

    library(tidyr)
    
    df1 %>% 
      gather(analysis, value, -year, -type, -type_group) %>%
      group_by(year, type, analysis) %>%
      summarise( value = diff(value)) %>%
      spread(analysis, value)
    

    【讨论】:

    • 谢谢,这很好用。当我有第二个问题时,我会查看我在问题中的错误并进行修改。只是最后一个问题 - 在代码的摘要部分,它采用下面的行并从上面的行中减去。有没有办法扭转这种情况,所以上面的行从下面的行中减去?或者我需要在应用上述代码之前对数据中的行进行重新排序吗?
    • 我认为最简单的方法是制作type = factor(type, levels = c(...)),然后根据您喜欢的顺序按type 排列。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2016-10-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-04-29
    • 2021-08-21
    相关资源
    最近更新 更多