【问题标题】:R: how to resample intraday data at the group level?R:如何在集团层面对日内数据进行重采样?
【发布时间】:2017-02-16 11:47:27
【问题描述】:

考虑以下数据框

time <-c('2016-04-13 23:07:45','2016-04-13 23:07:50','2016-04-13 23:08:45','2016-04-13 23:08:45'
         ,'2016-04-13 23:08:45','2016-04-13 23:07:50','2016-04-13 23:07:51')
group <-c('A','A','A','B','B','B','B')
value<- c(5,10,2,2,NA,1,4)
df<-data.frame(time,group,value)

> df
                 time group value
1 2016-04-13 23:07:45     A     5
2 2016-04-13 23:07:50     A    10
3 2016-04-13 23:08:45     A     2
4 2016-04-13 23:08:45     B     2
5 2016-04-13 23:08:45     B    NA
6 2016-04-13 23:07:50     B     1
7 2016-04-13 23:07:51     B     4

我想在5 seconds level - group level 处重新采样此数据帧,并计算每个 value 的 sum time-interval - @987654326 @。

区间左闭右开。例如,输出的第一行应该是

2016-04-13 23:07:45 A 5 因为第一个 5 秒间隔是 [2016-04-13 23:07:45, 2016-04-13 23:07:50[

如何在dplyr 或data.table 中做到这一点?我需要为时间戳导入lubridate 吗?

【问题讨论】:

  • 错字已修复。非常感谢!
  • 我认为来自data.table 的foverlaps 在这种情况下可能有用。我看看我能不能烤一个asnwer
  • 谢谢,或者用 dplyr group_by ?我不知道
  • @Noobie,只是为了澄清,我猜这个区间在右边是开放的......(这就是为什么你没有在第一行的总和中包含 23:07:50 值输出)?另外,如果组不同但发生在相同的时间间隔内,你会怎么处理(例如7 2016-04-13 23:07:51 B 4,假设下一行是8 2016-04-13 23:07:53 A 12)?我们是否忽略了组间的差异并简单地将 1、4 和 12 相加?
  • 嗨,约瑟夫,是的,左侧关闭,右侧打开。对于第二点,聚合在时间组级别,因此两个不同组的两个确切时间戳永远不会混淆在一起

标签: r data.table dplyr lubridate


【解决方案1】:

使用data.table的最新版本(1.9.8+):

library(data.table)

# convert to data.table, fix time, add future time
setDT(df)
df[, time := as.POSIXct(time)][, time.5s := time + 5]

# use non-equi join to filter on the required intervals and sum
df[, newval := df[df, on = .(group, time < time.5s, time >= time),
                  sum(value, na.rm = T), by = .EACHI]$V1]
df
#                  time group value             time.5s newval
#1: 2016-04-13 23:07:45     A     5 2016-04-13 23:07:50      5
#2: 2016-04-13 23:07:50     A    10 2016-04-13 23:07:55     10
#3: 2016-04-13 23:08:45     A     2 2016-04-13 23:08:50      2
#4: 2016-04-13 23:08:45     B     2 2016-04-13 23:08:50      2
#5: 2016-04-13 23:08:45     B    NA 2016-04-13 23:08:50      2
#6: 2016-04-13 23:07:50     B     1 2016-04-13 23:07:55      5
#7: 2016-04-13 23:07:51     B     4 2016-04-13 23:07:56      4

【讨论】:

  • 只需根据需要将on 中的不等式从严格更改为不严格。
  • 不应该把6和7放在一起吗?它们处于相同的 5 秒间隔内。看起来您只是为每个结果添加 5 秒。
  • 真的。 6和7应该加在一起。不错的收获@JosephWood
  • #7 在 #6 的未来区间内,但 #6 不在 #7 的未来区间内。如果您正在寻找未来+过去的匹配,请添加 time.neg.5s := time - 5 列并与那
  • 我感觉 time 和 time.5s 可以在同一个调用中创建,但我不确定这是否会提高速度
【解决方案2】:

这个怎么样:

library(dplyr)
Group5 <- function(myDf) {
    myDf$time <- ymd_hms(myDf$time)
    myDf$timeGroup <- floor_date(myDf$time, unit = "5 seconds")
    summarise(myDf %>% group_by(group, timeGroup), sum(value, na.rm = TRUE))
}

Group5(df)
Source: local data frame [5 x 3]
Groups: group [?]

   group           timeGroup `sum(value, na.rm = TRUE)`
  <fctr>              <dttm>                      <dbl>
1      A 2016-04-13 23:07:45                          5
2      A 2016-04-13 23:07:50                         10
3      A 2016-04-13 23:08:45                          2
4      B 2016-04-13 23:07:50                          5
5      B 2016-04-13 23:08:45                          2

它利用lubridate 中的floor_date 和ymd_hms 将每个日期时间放入适当的分组时间。

这是一个更奇特的例子:

set.seed(500)
time <- ymd_hms('2016-04-13 23:07:45') + sample(-10^3:10^3, 10^5, replace=TRUE)
group <- rep(LETTERS[1:20], each = 5000)
value <- rep(NA, 10^5)
value[sample(10^5, 95000)] <- sample(100, 95000, replace=TRUE)
df2 <- data.frame(time,group,value)

head(df2)
                 time group value
1 2016-04-13 23:18:53     A    53
2 2016-04-13 23:15:15     A    NA
3 2016-04-13 23:23:36     A    40
4 2016-04-13 23:06:40     A    23
5 2016-04-13 23:18:10     A    74
6 2016-04-13 22:57:56     A    65

我们有这样的称呼:

Group5(df2)
Source: local data frame [8,020 x 3]
Groups: group [?]

    group           timeGroup `sum(value, na.rm = TRUE)`
   <fctr>              <dttm>                      <int>
1       A 2016-04-13 22:51:05                        379
2       A 2016-04-13 22:51:10                        646
3       A 2016-04-13 22:51:15                        391
4       A 2016-04-13 22:51:20                       1118
5       A 2016-04-13 22:51:25                        745
6       A 2016-04-13 22:51:30                        546
7       A 2016-04-13 22:51:35                        884
8       A 2016-04-13 22:51:40                        711
9       A 2016-04-13 22:51:45                        526
10      A 2016-04-13 22:51:50                        484
# ... with 8,010 more rows

【讨论】:

  • 从 Op 的问题中得到了很好的猜测,但并没有考虑在 5 秒间隔上四舍五入。
【解决方案3】:

如果您愿意为每个组拥有单独的数据对象,您可以使用xts 来解决您的问题,而不是使用data.table,每个组对象。 xts period.apply 将自动处理您的区间在左侧关闭但在右侧也打开(这对于将金融报价数据汇总到条形频率非常有帮助。您不会在连续条形/间隔的区间边缘上重复计算报价):

time <-c('2016-04-13 23:07:45','2016-04-13 23:07:55','2016-04-13 23:08:45','2016-04-13 23:08:45'
         ,'2016-04-13 23:08:45','2016-04-13 23:07:50','2016-04-13 23:07:51')
group <-c('A','A','A','B','B','B','B')

value<- c(5,10,2,2,NA,1,4)
df=data.frame(time,group,value)

library(quantmod)
library(lubridate)
df$time = ymd_hms(df$time)

# In this example, model group B object: (You can easily generalise this with a loop or lapply over multiple groups)
df_grp <- df[df$group == "B", ]
x.df_grp <- xts(df_grp$value, order.by = df_grp$time) 
ep <- endpoints(x.df_grp, on = "seconds", k = 5)
# You can replace sum by any useful function.  Pass in extra arguments to period.apply that correspond to FUN, here na.rm = T, to avoid having sum returning NA in your group B row:
x.df_grp_5sec <- period.apply(x.df_grp, ep, FUN = sum, na.rm = TRUE)
# Align timestamps to end of each 5 sec interval by default (helps avoid lookforward bias when merging time series data on different time frequencies):
x.df_grp_5sec <- align.time(x.df_grp_5sec, 5)
# Now record timestamps at start of each 5 sec interval:
.index(x.df_grp_5sec) <- .index(x.df_grp_5sec) - 5

#result:
> x.df_grp_5sec
                    [,1]
2016-04-13 23:07:50    5
2016-04-13 23:08:45    2

【讨论】:

    【解决方案4】:

    我想到data.table:

    library(data.table)
    setDT(df)
    df[, result:={lv=df$group==group; dt=difftime( df$time, time, units="sec"); print(dt); sum(df$value[lv & dt >= 0 & dt < 5],na.rm=TRUE)},by=1:nrow(df)]
    

    输出:

                      time group value result
    1: 2016-04-13 23:07:45     A     5      5
    2: 2016-04-13 23:07:50     A    10     10
    3: 2016-04-13 23:08:45     A     2      2
    4: 2016-04-13 23:08:45     B     2      2
    5: 2016-04-13 23:08:45     B    NA      2
    6: 2016-04-13 23:07:50     B     1      5
    7: 2016-04-13 23:07:51     B     4      4
    

    j 部分详情:

    lv=df$group==group # Create a logical vector to filter at end
    dt=abs( difftime( df$time, time, units="sec")) # compute the time difference in seconds between current row and all others
     sum(df$value[lv & dt >= 0 & dt < 5]) # Sum the values where in same group and the difference in seconds is between 0 and 5 secs, 0 included, 5 excluded 
    

    result:={} 允许我们将结果创建为函数调用。 by=1:nrow(df) 让它逐行工作。

    并过滤结果以仅获取起点:

    > df[,.SD[!duplicated(result)],by=group]
       group                time value result
    1:     A 2016-04-13 23:07:45     5      5
    2:     A 2016-04-13 23:07:50    10     10
    3:     A 2016-04-13 23:08:45     2      2
    4:     B 2016-04-13 23:08:45     2      2
    5:     B 2016-04-13 23:07:50     1      5
    6:     B 2016-04-13 23:07:51     4      4
    

    【讨论】:

    • 为什么第 7 行的 result 等于 5?上一行不在[time, time+5) 范围内。
    • @eddi 第 6 行在 5 秒内的范围内,如果不需要,只需删除 difftime 左右的 abs 调用
    • @Eddie,我从 Op 的 Q 中推断出范围,确实在这种情况下,abs 调用对于答案来说是多余的
    • @Tensibai,不确定这是一个问题还是建议......我们是否应该在求和时删除 NA?
    • 另外,您需要将data.frame 更改为data.table.. 在我意识到我需要data.table 之前,我花了几次错误。
    猜你喜欢
    • 2015-01-31
    • 2011-10-13
    • 2018-10-29
    • 2021-01-25
    • 2016-09-03
    • 1970-01-01
    • 2017-05-20
    • 1970-01-01
    • 2018-11-22
    相关资源
    最近更新 更多