【发布时间】:2018-01-14 16:15:50
【问题描述】:
我有以下数据集,我想对每个序列进行分组和总结。每个序列应拆分为所有事件,这些事件发生在第一个日期后的前 7 天,并将后面的事件合并为一个单独的组。基本上我最大的挑战是找到序列中的第一个日期,添加 7 天并标记该序列中属于该类别的所有日期。
structure(list(`Sequence ID` = c("1_0_0", "1_0_0", "1_0_0", "1_0_0",
"1_0_0", "1_1_0", "1_1_0", "1_1_0", "1_1_0", "1_1_0", "1_2_0",
"1_2_1", "1_2_1", "1_2_1", "1_2_1", "1_2_2"), Date = c("02.12.2015 20:16",
"03.12.2015 20:17", "02.12.2015 20:44", "03.12.2015 09:32", "03.12.2015 09:33",
"07.12.2015 08:18", "08.12.2015 19:40", "08.12.2015 19:43", "22.12.2015 18:22",
"22.12.2015 18:23", "23.12.2015 14:18", "05.01.2016 11:35", "05.01.2016 13:21",
"05.01.2016 13:22", "05.01.2016 13:22", "04.08.2016 08:25"),
StimuliA = c(0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 1L,
0L, 0L, 0L, 0L, 0L), StimuliB = c(0L, 0L, 0L, 0L, 0L, 0L,
0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 1L), Response = c(1L,
1L, 1L, 1L, 1L, 0L, 1L, 1L, 1L, 1L, 0L, 0L, 1L, 1L, 1L, 0L
)), .Names = c("Sequence ID", "Date", "StimuliA", "StimuliB",
"Response"), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA,
-16L), spec = structure(list(cols = structure(list(`Sequence ID` = structure(list(), class = c("collector_character",
"collector")), Date = structure(list(), class = c("collector_character",
"collector")), StimuliA = structure(list(), class = c("collector_integer",
"collector")), StimuliB = structure(list(), class = c("collector_integer",
"collector")), Response = structure(list(), class = c("collector_integer",
"collector")), X6 = structure(list(), class = c("collector_skip",
"collector")), X7 = structure(list(), class = c("collector_skip",
"collector")), X8 = structure(list(), class = c("collector_skip",
"collector")), X9 = structure(list(), class = c("collector_skip",
"collector")), X10 = structure(list(), class = c("collector_skip",
"collector"))), .Names = c("Sequence ID", "Date", "StimuliA",
"StimuliB", "Response", "X6", "X7", "X8", "X9", "X10")), default = structure(list(), class = c("collector_guess",
"collector"))), .Names = c("cols", "default"), class = "col_spec"))
这可能是一个可能的输出,其中 Group 0 总结了前 7 天的所有值,1 总结了后来发生的值。
Sequence ID Group Date StimuliA StimuliB Response
1_0_0 0 02.12.2015 20:16 0 0 5
1_0_0 1 09.12.2015 20:16 0 0 0
1_1_0 0 07.12.2015 08:18 1 0 2
1_1_0 1 14.12.2015 08:18 0 0 2
1_2_0 0 23.12.2015 14:18 1 0 0
1_2_0 1 30.12.2015 14:18 0 0 0
1_2_1 0 05.01.2016 11:35 0 1 3
1_2_1 1 12.01.2016 11:35 0 0 0
1_2_2 0 04.08.2016 08:25 0 1 0
1_2_2 1 11.08.2016 08:25 0 0 0
我会尝试使用以下代码来实现这一点,但需要一些输入如何识别 7 天前后的值。
#change the date into posixct format
df$Date <- as.POSIXct(strptime(master$Date,"%d.%m.%Y %H:%M"))
#arrange the dataframe according to User and Date
df <- arrange(df, Sequence ID,Date)
#identify the values before and after 7 days
#aggregate all the eventlog rows according to the stimuli IDs
df <- aggregate(. ~ Sequence ID + Group, data=df, sum)
【问题讨论】: