【问题标题】:Merging cases into one in R在R中将案例合并为一个
【发布时间】:2013-11-30 17:55:41
【问题描述】:

我有一个非常新手的问题。我正在使用 Aid Worker Security Database,该数据库记录了针对援助人员的暴力事件,以及从 1997 年至今的事件报告。事件在数据集中独立标记。我想合并给定年份在一个国家发生的所有事件,将其他变量的值相加,并为所有国家(1997-2013)创建一个具有相同年数的简单时间序列。知道怎么做吗?

df
#   year  country totalnationals internationalskilled
# 1 1997   Rwanda              0                    3
# 2 1997 Cambodia              1                    0
# 3 1997  Somalia              0                    1
# 4 1997   Rwanda              1                    0
# 5 1997 DR Congo             10                    0
# 6 1997  Somalia              1                    0
# 7 1997   Rwanda              1                    0
# 8 1998   Angola              5                    0

其中“df”定义为:

df <- structure(list(year = c(1997L, 1997L, 1997L, 1997L, 1997L, 1997L, 
  1997L, 1998L), country = c("Rwanda", "Cambodia", "Somalia", "Rwanda", 
  "DR Congo", "Somalia", "Rwanda", "Angola"), totalnationals = c(0L, 
  1L, 0L, 1L, 10L, 1L, 1L, 5L), internationalskilled = c(3L, 0L, 
  1L, 0L, 0L, 0L, 0L, 0L)), .Names = c("year", "country", "totalnationals", 
  "internationalskilled"), class = "data.frame", row.names = c(NA, -8L))

我想要这样的东西:

#    year  country totalnationals internationalskilled
# 1  1997   Rwanda              2                    3
# 2  1997 Cambodia              1                    0
# 3  1997  Somalia              1                    1
# 4  1997 DR Congo             10                    0
# 5  1997   Angola              0                    0
# 6  1998   Rwanda              0                    0
# 7  1998 Cambodia              0                    0
# 8  1998  Somalia              0                    0
# 9  1998 DR Congo              0                    0
# 10 1998   Angola              5                    0

很抱歉这个非常非常新手的问题......但到目前为止我不知道该怎么做。谢谢! :-)

【问题讨论】:

  • 请阅读this,然后相应地编辑您的问题。

标签: r reshape


【解决方案1】:

在 OP 的 cmets 之后更新 -

df <- subset(df, year <= 2013 & year >= 1997)
df$totalnationals <- as.integer(df$totalnationals)
df$internationalskilled <- as.integer(df$internationalskilled)
df2 <- aggregate(data = df,cbind(totalnationals,internationalskilled)~year+country, sum)

为没有记录的年份添加 0 -

df3 <- expand.grid(unique(df$year),unique(df$country))
df3 <- merge(df3,df2, all.x = TRUE, by = 1:2)
df3[is.na(df3)] <- 0

【讨论】:

  • 感谢您的回答,但没有达到我的预期。它要么显示“评估错误(expr,envir,enclos):找不到对象'year'',或者如果我包含'df = 1997)'。无论如何谢谢:)
  • 已更新。此外,我可能会遗漏一个方面,即如果你需要一个 0 多年而没有人被杀。这是必需的吗?
  • 也添加了插入零部分。
  • 零部分的代码有效,但我最终得到了相同国家和年份的重复值。合并后,我最终得到了大约 150 万个案例,我应该有大约 1190 个案例(17 年(1997-2013)x 70 个国家/地区)。有没有办法消除重复的情况?谢谢!
  • 对不起,我的错。 expand.grid 需要一个 unique 在里面。现在可以试试了吗?
【解决方案2】:

数据表也是如此(在大型数据集上可能更快)。

library(data.table)
dt   <- data.table(df,key="year,country")
smry <- dt[,list(totalnationals      =sum(totalnationals), 
                 internationalskilled=sum(internationalskilled)),
           by="year,country"]
countries   <- unique(dt$country)
template    <- data.table(year=rep(1997:2013,each=length(countries)),
                          country=countries, 
                          key="year,country")
time.series <- smry[template]
time.series[is.na(time.series)]=0

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-12-14
    • 2019-08-13
    • 1970-01-01
    • 1970-01-01
    • 2020-07-28
    • 1970-01-01
    相关资源
    最近更新 更多