【问题标题】:Summing consecutive values, broken up by specific value, in R对 R 中的连续值求和,按特定值分解
【发布时间】:2021-08-03 05:28:18
【问题描述】:

我无法弄清楚如何对变量进行分组以从 dplyr 获得所需的结果。我有一个这样的实验数据集:

  subject   task_phase  block_number trial_number ResponseCorrect
   <chr>     <chr>              <dbl>        <dbl>           <dbl>
 1 268301377    1            1            2               1
 2 268301377    1            1            3               1
 3 268301377    1            1            4               1
 4 268301377    1            2            2              -1
 5 268301377    1            2            3               1
 6 268301377    1            2            4               1
 7 268301377    1            3            2               1
 8 268301377    1            3            3              -1
 9 268301377    1            3            4               1
10 268301377    2            1           50               1
11 268301377    2            1           51               1
12 268301377    2            1           52               1
13 268301377    2            2           37              -1
14 268301377    2            2           38               1
15 268301377    2            2           39               1
16 268301377    2            3           41              -1
17 268301377    2            3           42              -1
18 268301377    2            3           43               1

我希望对连续的“正确”响应求和,并在每次出现错误响应时“重置”这个计数:

  subject   task_phase  block_number trial_number ResponseCorrect   ConsecutiveCorrect
   <chr>     <chr>              <dbl>        <dbl>       <dbl>            <dbl>
 1 268301377    1            1            1               1                 1
 2 268301377    1            1            2               1                 2
 3 268301377    1            1            3               1                 3
 4 268301377    1            2            1              -1                 0
 5 268301377    1            2            2               1                 1
 6 268301377    1            2            3               1                 2
 7 268301377    1            3            1               1                 1
 8 268301377    1            3            2              -1                 0
 9 268301377    1            3            3               1                 1
10 268301377    2            1            1               1                 1
11 268301377    2            1            2               1                 2
12 268301377    2            1            3               1                 3
13 268301377    2            2            1              -1                 0
14 268301377    2            2            2               1                 1
15 268301377    2            2            3               1                 2
16 268301377    2            3            1              -1                 0
17 268301377    2            3            2              -1                 0
18 268301377    2            3            3               1                 1

我最初认为我可以按照df %&gt;% group_by(subject, task_phase, block_number, ResponseCorrect) %&gt;% mutate(ConsecutiveCorrect = cumsum(ResponseCorrect) 的方式做一些事情,并且几乎 有效。但是,它没有给出连续的值:它只是总结了每个块的正确响应总数(。我实际上是在尝试使用 -1s 作为重新开始求和的断点。

是否有我不知道的分组功能(Tidyverse 或其他)可以按照这些方式做一些事情?

【问题讨论】:

  • 这两个答案都解决了手头的问题 - 感谢你们!

标签: r dplyr tidyverse


【解决方案1】:

你可以试试

library(dplyr)

data %>% 
  group_by(
    subject, 
    task_phase, 
    block_number, 
    grp = lag(cumsum(ResponseCorrect == -1), default = 0)
    ) %>% 
  mutate(ConsecutiveCorrect = ifelse(ResponseCorrect == -1, 0, cumsum(ResponseCorrect))) %>% 
  ungroup() %>% 
  select(-grp)

返回

# A tibble: 18 x 6
     subject task_phase block_number trial_number ResponseCorrect ConsecutiveCorrect
       <dbl>      <dbl>        <dbl>        <dbl>           <dbl>              <dbl>
 1 268301377          1            1            2               1                  1
 2 268301377          1            1            3               1                  2
 3 268301377          1            1            4               1                  3
 4 268301377          1            2            2              -1                  0
 5 268301377          1            2            3               1                  1
 6 268301377          1            2            4               1                  2
 7 268301377          1            3            2               1                  1
 8 268301377          1            3            3              -1                  0
 9 268301377          1            3            4               1                  1
10 268301377          2            1           50               1                  1
11 268301377          2            1           51               1                  2
12 268301377          2            1           52               1                  3
13 268301377          2            2           37              -1                  0
14 268301377          2            2           38               1                  1
15 268301377          2            2           39               1                  2
16 268301377          2            3           41              -1                  0
17 268301377          2            3           42              -1                  0
18 268301377          2            3           43               1                  1

【讨论】:

    【解决方案2】:

    data.table 的选项。按'subject'、'task_phase'、'block_number'分组,得到'ResponseCorrect'的run-length-id (rleid),返回该序列的rowid,乘以一个逻辑向量,使得对应的元素为 -1(FALSE -> 0 将返回 0,TRUE -> 1 返回元素)

    library(data.table)
    setDT(df)[, ConsecutiveCorrect := rowid(rleid(ResponseCorrect)) *
             (ResponseCorrect == 1), by = .(subject, task_phase, block_number)]
    

    -输出

    df
          subject task_phase block_number trial_number ResponseCorrect ConsecutiveCorrect
     1: 268301377          1            1            2               1                  1
     2: 268301377          1            1            3               1                  2
     3: 268301377          1            1            4               1                  3
     4: 268301377          1            2            2              -1                  0
     5: 268301377          1            2            3               1                  1
     6: 268301377          1            2            4               1                  2
     7: 268301377          1            3            2               1                  1
     8: 268301377          1            3            3              -1                  0
     9: 268301377          1            3            4               1                  1
    10: 268301377          2            1           50               1                  1
    11: 268301377          2            1           51               1                  2
    12: 268301377          2            1           52               1                  3
    13: 268301377          2            2           37              -1                  0
    14: 268301377          2            2           38               1                  1
    15: 268301377          2            2           39               1                  2
    16: 268301377          2            3           41              -1                  0
    17: 268301377          2            3           42              -1                  0
    18: 268301377          2            3           43               1                  1
    

    数据

    df <- structure(list(subject = c(268301377L, 268301377L, 268301377L, 
    268301377L, 268301377L, 268301377L, 268301377L, 268301377L, 268301377L, 
    268301377L, 268301377L, 268301377L, 268301377L, 268301377L, 268301377L, 
    268301377L, 268301377L, 268301377L), task_phase = c(1L, 1L, 1L, 
    1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), 
        block_number = c(1L, 1L, 1L, 2L, 2L, 2L, 3L, 3L, 3L, 1L, 
        1L, 1L, 2L, 2L, 2L, 3L, 3L, 3L), trial_number = c(2L, 3L, 
        4L, 2L, 3L, 4L, 2L, 3L, 4L, 50L, 51L, 52L, 37L, 38L, 39L, 
        41L, 42L, 43L), ResponseCorrect = c(1L, 1L, 1L, -1L, 1L, 
        1L, 1L, -1L, 1L, 1L, 1L, 1L, -1L, 1L, 1L, -1L, -1L, 1L)), 
    class = "data.frame", row.names = c("1", 
    "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", 
    "14", "15", "16", "17", "18"))
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2022-11-15
      • 1970-01-01
      • 1970-01-01
      • 2020-04-08
      • 1970-01-01
      • 2021-11-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多