【发布时间】:2014-11-03 22:27:45
【问题描述】:
我正在尝试确定两次观察之间的时间差。数据由不同的个人分解,每个人都有自己的唯一 ID。我有一个数据集,它告诉我每次更改时他们的状态更新,以及他们的状态何时更改。状态可以是两个值之一,并且它总是会更改为不是的值(在这种情况下,从 Y 到 N,或从 N 到 Y)。
数据如下:
ID Status Time
1 Y 2013-07-01 08:07:00
2 Y 2013-07-01 08:07:03
3 Y 2013-07-01 08:07:04
4 Y 2013-07-01 08:07:06
1 N 2013-07-01 08:07:07
2 N 2013-07-01 08:07:23
5 Y 2013-07-01 08:07:34
6 Y 2013-07-01 08:07:45
7 Y 2013-07-01 08:07:47
1 Y 2013-07-01 08:07:56
3 N 2013-07-01 08:07:58
我想找到的是每个 ID 的每次状态更改之间经过的时间量 - 即从 Y 到 N 需要多长时间。然后获得汇总统计信息,例如经过的分布时间,经过时间的平均值等。
因此示例输出可能如下所示,记录上面发生的三个 Y 到 N 切换(1 个切换、2 个切换和 3 个切换)
Y to N change Time elapsed (in seconds)
1 7
2 20
3 54
由于某种原因,我遇到了很多麻烦。现在我有 POSIXlt 格式的时间,以及 ID 和状态作为一个因素。我曾尝试使用 ddply 按 ID 然后按时间戳对数据进行排序,但到目前为止还没有奏效。任何建议将不胜感激!
编辑:将时间更改为正确的类型。
Edit2:在等待更多答案时最终编写了一个解决方案。我的方式比这里的许多解决方案都要丑陋得多,但我做到了:
N <- ifelse(df$Status=="N",1,0)
Y <- ifelse(df$Status== "Y",1,0)
#making a vector which is 1 for a row if the item status of the row below it is N
var1 <- N
for (i in 1:nrow(df)) {
var1[i] <- N[i+1]
}
#making a vector which is TRUE if a row's item status is Y and the row after is N
check <- ifelse(var1==s & var1==1,TRUE,FALSE)
#had to define the last one as FALSE manually because the for loop above would miss the last entry due to how it was constructed
check [50000]=FALSE
#made a loop which finds the time difference for a row's TIME and the row below it, given that "check " is true for that row, and writes that to a results vector.
#here is the results vector
results <- numeric(nrow(df))
#here is the for loop
for (i in 1:nrow(df)) {
if(check [i]){
results[i] <- difftime(df$Time[i],df$Time[i+1])
}
}
我最初使用 for 循环解决了这个问题,但是在我实际数据集的大约 100 万行中,它太慢了,所以我做了这个矢量化的东西。这些其他解决方案是否适用于这么大的数据?我一定会尝试一下!
【问题讨论】: