【问题标题】:increment column records based on changes in other column in R根据R中其他列的变化增加列记录
【发布时间】:2020-09-25 18:11:36
【问题描述】:

当同一会话的第一个 timestamp 与后续 timestamp 记录之间的差异超过 10 个单位时,我想在 session 列中添加 1。

换句话说:
如果同一会话中timestamp 列中的间隙大于 10,则为特定 ID 的其余会话添加 1。所以我们不应该有相同的session,其记录中的差距超过10。
让我们说:

df<-read.table(text="
ID      timestamp    session
1       10             1
1       12             1
1       15             1
1       21             1
1       25             1
1       27             2
1       29             2
2       11             1
2       22             2
2       27             2
2       32             2
2       42             2
2       43             3",header=T,stringsAsFactors = F)

在上面的示例中,对于ID==1,第 4 行中的会话间隔 (timestamp==10) 大于 10 (timestamp==21),因此我们将其余会话加 1。每当会话数发生变化时,同一个会话中时间戳的第一条记录的差值应小于10,否则应添加到会话中。

result:  

ID      timestamp    session
1      *10             1
1       12             1
1       15             1
1      *21             2     <-- because 21-10 >= 10 it add 1 to the rest of sessions in this ID 
1       25             2
1       27             3
1       29             3
2       11             1
2      *22             2
2       27             2
2      *32             3     <-- because 32-22>= 10 it add 1 to the rest of session
2      *42             4     <-- because 42-32>=10
2       43             5

如何在 R 中做到这一点?

【问题讨论】:

  • 与第一个戳记的差值大于10的所有值是否必须加1?
  • 第二组第一个时间戳也是11,为什么用22?
  • @Duck 是的,与第一个值的区别是它们具有相同的会话数。关于你的第二个问题,它仍然是 11。22 是第二个时间戳。

标签: r


【解决方案1】:

也许自定义函数可能有助于计算累积和并在达到阈值后重置。在这种情况下,如果您为函数提供 session 数据,它将提供一个包含会话累积“偏移量”的结果,但仅在会话数未增加时以行的形式提供。这解决了 ID 2 timestamp 22 的情况,其中差异 > 10,但会话数从 1 增加到 2。

library(tidyverse)

threshold <- 10

cumsum_with_reset <- function(x, session, threshold) {
  cumsum <- 0
  group <- 0
  result <- numeric()
  for (i in seq_along(x)) {
    cumsum <- cumsum + x[i]
    if (cumsum >= threshold) {
      if (session[i] == session[i-1]) {
        group <- group + 1
      }
      cumsum <- 0
    }
    result = c(result, group)
  }
  return (result)
}

df %>%
  group_by(ID) %>%
  mutate(diff = c(0, diff(timestamp)),
         cumdiff = cumsum_with_reset(diff, session, threshold),
         new_session = cumdiff + session)

函数改编自this solution。

输出

      ID timestamp session  diff cumdiff new_session
   <int>     <int>   <int> <dbl>   <dbl>       <dbl>
 1     1        10       1     0       0           1
 2     1        12       1     2       0           1
 3     1        15       1     3       0           1
 4     1        21       1     6       1           2
 5     1        25       1     4       1           2
 6     1        27       2     2       1           3
 7     1        29       2     2       1           3
 8     2        11       1     0       0           1
 9     2        22       2    11       0           2
10     2        27       2     5       0           2
11     2        32       2     5       1           3
12     2        42       2    10       2           4
13     2        43       3     1       2           5

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-10-10
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多