【发布时间】:2013-06-05 20:05:42
【问题描述】:
我正在使用 R 语言处理一个包含 3 个因素的大型数据集:FY(6 个级别)、Region(10 个级别)和 Service(24 个级别)。我需要在所有三个级别上对我的数字向量 SumOfUnits 求和,我能想到的唯一方法是将数据帧首先拆分为:6 个数据帧,按 FY 拆分,然后将这 6 个数据帧拆分为 10 个数据帧,按区域拆分,然后将这 10 个数据帧分成 24 个服务,然后我终于可以将数字向量的总和重新组合成一个数据帧。该数据框将具有 6*10*24 (1440) 行和 4 列。我目前这样做的方式涉及很多拆分,所以我认为可能有一个我可以编写的函数,我可以在拆分的每个级别使用,但我在 R 中没有使用太多“函数”所以我不确定要写什么(如果有的话)。我还想可能有一种更有效的方法来获取格式化的数据集,所以我欢迎所有建议。
以下是我的数据框中的几行:
FY Region Service SumOfUnits
1 2006 1 Medication 13
2 2006 1 Medication 1
3 2006 1 Screening & Assessment 38
4 2006 1 Screening & Assessment 13
5 2006 1 Screening & Assessment 41
6 2006 1 Screening & Assessment 67
7 2006 1 Screening & Assessment 222
8 2006 1 Residential Treatment 38
9 2006 1 Residential Treatment 1558
这是我一直用于拆分的代码:
# Creating a data frame by year
X <- split(MIC, MIC$FY)
Y <- lapply(seq_along(X), function(x) as.data.frame(X[[x]])[, ])
#Assign the dataframes in the list Y to individual objects
A <- Y[[1]]
B <- Y[[2]]
C <- Y[[3]]
D <- Y[[4]]
E <- Y[[5]]
Q <- Y[[6]]
#Creating 10 dataframes from 2006 split by region
X <- split(A, A$Region)
Y <- lapply(seq_along(X), function(x) as.data.frame(X[[x]])[, ])
Reg1 <- Y[[1]]
Reg2 <- Y[[2]]
Reg3<- Y[[3]]
Reg4 <- Y[[4]]
Reg5<- Y[[5]]
Reg6 <- Y[[6]]
Reg7 <- Y[[7]]
Reg8 <- Y[[8]]
Reg9 <- Y[[9]]
Reg10<- Y[[10]]
#Creating 24 dataframes: for 2006, region 1
X <- split(Reg1, Reg1$Service)
Y <- lapply(seq_along(X), function(x) as.data.frame(X[[x]])[, ])
Serv1 <- Y[[1]]
Serv2 <- Y[[2]]
Serv3<- Y[[3]]
Serv4 <- Y[[4]]
Serv5<- Y[[5]]
#etc...
我希望我的数据样本看起来像这样:
FY Region Service SumOfUnits
2006 1 Medication 4300
2006 2 Medication 3299
2006 3 Medication 2198
2007 1 Medication 5467
2007 2 Medication 3214
2007 3 Medication 9807
【问题讨论】:
-
你看过 plyr 吗?
-
...甚至只是
aggregate,对吧? -
我都试过了,但不知道如何在 3 个因素的水平上使用它们。因此,例如,我可以执行以下操作:
library(plyr) Sum_Year <- na.omit(ddply(MIC[c(4)], .(FY, Region, Service), colSums,na.rm=TRUE))但由于每个服务的行数不均匀,我会收到错误消息。如果我只使用 FY 运行相同的代码,它可以工作,但我需要按 FY、区域和服务分解它 -
对于聚合,你会做类似
aggregate(SumOfUnits ~ FY + Region + Service,data = MIC,FUN = sum)的事情。