这是一种替代解决方案,它使用foverlaps() 将给定的时间范围分成一天长度的片段,并为每个片段计算process_duration。
library(data.table)
library(lubridate)
# create vector of start dates
start_date <- setDT(df)[, seq(floor_date(min(start_date_time), "day"),
max(end_date_time),
by = "1 day")]
# create keyed data.table with start and end of each day
day_grid <- data.table(start_date,
end = start_date + days(1),
key = "start_date,end")
# find overlaps of ranges in df with day_grid
df2 <- foverlaps(df, day_grid, by.x = c("start_date_time", "end_date_time"))
# compute durations
df2[, process_duration := difftime(
pmin(end, end_date_time),
pmax(start_date, start_date_time),
units = "hours")][
# clean up
process_duration > 0, .(start_date, process_duration, random_col)][
# sort output
order(start_date)]
start_date process_duration random_col
1: 2019-01-01 18.37806 hours blabla
2: 2019-01-01 12.00000 hours dddd
3: 2019-01-02 10.40667 hours blabla
4: 2019-01-02 20.00000 hours dddd
5: 2019-01-03 24.00000 hours dddd
6: 2019-01-04 12.00000 hours dddd
7: 2019-01-05 24.00000 hours eeee
8: 2019-01-06 24.00000 hours eeee
9: 2019-01-07 2.00000 hours eeee
这种方法的优点是可以轻松适应不同的时间网格,例如小时、周或月。
difftime 对象具有units 属性。因此,列名缩写为process_duration。
数据
为了比较,我们使用了arg0naut's answer 的增强数据集。字符日期时间立即被 ymd_hms() 强制转换为 POSIXct。
df <- data.frame(
start_date_time = ymd_hms(c(
"2019-01-01 05:37:19",
"2019-01-01 03:15:01",
"2019-01-02 04:00:00",
"2019-01-05 00:00:00"
)),
process_duration_in_hours = c(28.78, 12.00, 56.00, 50.00),
end_date_time = ymd_hms(c(
"2019-01-02 10:24:24",
"2019-01-01 15:15:01",
"2019-01-04 12:00:00",
"2019-01-07 02:00:00"
)),
random_col = c("blabla", "dddd", "dddd", "eeee")
)