【问题标题】:Replace column values based on column in another dataframe根据另一个数据框中的列替换列值
【发布时间】:2019-12-02 08:10:03
【问题描述】:

我想根据另一个数据框中的列替换 df 中的一些列值 这是第一个df的头:

 df1
A tibble: 253 x 2
      id sum_correct
    <int>       <dbl>
 1 866093          77
 2 866097          95
 3 866101          37
 4 866102          65
 5 866103          16
 6 866104          72
 7 866105          99
 8 866106          90
 9 866108          74
10 866109          92

有些 sum_correct 需要用另一个 df 中的正确值替换,使用 id 来触发替换

df 2 
A tibble: 14 x  2
     id sum_correct
    <int>       <dbl>
 1 866103          61
 2 866124          79
 3 866152          85
 4 867101          24
 5 867140          76
 6 867146          51
 7 867152          56
 8 867200          50
 9 867209          97
10 879657          56
11 879680          61
12 879683          58
13 879693          77
14 881451          57

如何在 R studio 中实现这一点?我在这里先向您的帮助表示感谢。

【问题讨论】:

    标签: r dataframe


    【解决方案1】:

    您可以使用match 进行更新加入,以查找id 匹配的位置,并使用which 删除不匹配的(NA):

    idx <- match(df1$id, df2$id)
    idxn <- which(!is.na(idx))
    df1$sum_correct[idxn]  <- df2$sum_correct[idx[idxn]]
    df1
           id sum_correct
    1  866093          77
    2  866097          95
    3  866101          37
    4  866102          65
    5  866103          61
    6  866104          72
    7  866105          99
    8  866106          90
    9  866108          74
    10 866109          92
    

    【讨论】:

      【解决方案2】:

      您可以使用left_join,然后使用coalesce:

      library(dplyr)
      left_join(df1, df2, by = "id", suffix = c("_1", "_2")) %>%
        mutate(sum_correct_final = coalesce(sum_correct_2, sum_correct_1))
      

      新列sum_correct_final 包含来自df2 的值(如果存在)和来自df1 的值(如果来自df2 的对应条目不存在)。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2020-09-01
        • 2021-06-11
        • 2019-01-31
        • 2022-07-30
        • 2021-06-29
        • 1970-01-01
        • 2019-09-02
        • 2021-11-12
        相关资源
        最近更新 更多