【问题标题】:R - Find and Update Based On Most Recent Matching ColumnsR - 根据最近的匹配列查找和更新
【发布时间】:2021-11-07 08:19:26
【问题描述】:

我有一个大数据集:

head(data)

  subject stim1 stim2 Chosen outcome
1       1     2     1      2       0
2       1     3     2      2       0
3       1     3     1      1       0
4       1     2     3      3       1
5       1     1     3      1       1
6       1     2     1      1       1

tail(data)
      subject stim1 stim2 Chosen outcome
44249    3020    40    42     42       0
44250    3020    40    41     41       1
44251    3020    44    45     45       1
44252    3020    41    43     43       0
44253    3020    42    40     42       0
44254    3020    42    44     44       1

我的目标是(在每个主题内)为每一行检查最近出现相同两个 stim1 和 stim2 的案例,然后添加一列

  1. 从该行选择的条目 (Previous_Choice)
  2. 该行的结果变量 (Previous_outcome)
  3. 之前未在该行中选择的数字是否(即在 Previous_Choice 行中) 随后在导致当前试验的任何行中被选中。例如,如果它的 stim1=1 和 stim2=2 和 Chosen=2,那么我正在查看在此之后的任何试验中是否 Chosen=1(直到我的当前行)(S_choice)(例如,参见第 6 行)李>

棘手的部分是我不关心哪个其中一个数字是stim1,哪个数字是stim2。 For example if my current trial stim1=1 and stim2=2 i want the most recent trial where (stim1=1,stim2=2 OR stim1=2, stim2=1)

期望的结果

  subject stim1 stim2 Chosen outcome   Previous_Choice  Previous_Outcome  S_choice 
1       1     2     1      2       0         NA                 NA         NA
2       1     3     2      2       0         NA                 NA         NA
3       1     3     1      1       0         NA                 NA         NA
4       1     2     3      3       1          2                 0        FALSE
5       1     1     3      1       1          1                 0        FALSE
6       1     2     1      1       1          2                 0        TRUE

注意 - S_choice 在第六行中为真的原因是因为在试验 1 之后(其中 1 和 2 分别是 stim1 和 stim2),在第 3 行和第 5 行中选择了 1

  str(data)
'data.frame':   44254 obs. of  5 variables:
 $ subject: num  1 1 1 1 1 1 1 1 1 1 ...
 $ stim1  : int  2 3 3 2 1 2 2 3 2 2 ...
 $ stim2  : int  1 2 1 3 3 1 3 1 1 1 ...
 $ Chosen : int  2 2 1 3 1 1 2 1 2 2 ...
 $ outcome: int  0 0 0 1 1 1 1 0 1 0 ...

【问题讨论】:

    标签: r database dataframe dplyr


    【解决方案1】:

    我不明白 S_choise 是什么意思,但也许我可以帮助您处理其他 2 列。

    LastOrNa <- function(x) {
      if (length(x) == 0) {
        return(NA)
      }
      return(last(x))
    }
    
    LastEq <- function(x, y) {
      res <- sapply(2:length(x), function(t) {
        LastOrNa(which(
            (x[1:(t - 1)] == x[t] & y[1:(t - 1)] == y[t]) |
             (x[1:(t - 1)] == y[t] & y[1:(t - 1)] == x[t])
          ))
        }
      )
      return(c(NA, res))
    }
    
    data %>% group_by(subject) %>% 
      mutate(
        last_eq = LastEq(stim1, stim2),
        Previous_Choice = Chosen[last_eq],
        Previous_Outcome = outcome[last_eq],
        last_eq = NULL
      )
    

    【讨论】:

    • 是的,您的答案非常适合前两列,谢谢!
    • 我试图编辑问题以使 S_Choice 更清晰 - 我也试图举一个例子来澄清它。
    • 哦。我的英语很差,对不起。我对 S_Choice 有一些疑问: 1) 在第六行中,您有 1 和 2 作为 stim1 和 stim2。 Chosen 等于 1。第 3 行和第 5 行在 Chosen 列中的编号为 1。你想找到相等的数字吗? 2) 为什么 S_Choice 在第五行是 FALSE? 3) 在描述中,您提到了有关 Previous_choice 列的内容,但我不明白您想如何使用它。
    • 您的解决方案实际上足以找到 S_choice,再次感谢您!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-11-26
    • 2018-03-01
    • 1970-01-01
    • 1970-01-01
    • 2015-01-23
    • 2021-02-08
    相关资源
    最近更新 更多