【问题标题】:How to recode values based on duplicate values in another dataset如何根据另一个数据集中的重复值重新编码值
【发布时间】:2019-11-25 02:27:30
【问题描述】:

我正在处理以下数据。我们可以叫它x

   New_Name_List               Old_Name_List
1     bumiputera        bumiputera (muslims)
2     bumiputera          bumiputera (other)
3 non bumiputera non bumiputera (indigenous)
4        chinese                     chinese

目标是在另一个看起来像这样的数据对象中重新编码数据。我们可以叫它y

  EPR_country_code EPR_country           EPR_group_lower_2
1              835      Brunei        bumiputera (muslims)
2              835      Brunei          bumiputera (other)
3              835      Brunei non bumiputera (indigenous)
4              835      Brunei                     chinese 

如果x$New_Name_List 有重复值,我希望新列y$EPR_group_lower_3 中的x$Old_Name_List 值。

如果x$New_Name_List 具有唯一值,我希望新列y$EPR_group_lower_3 中的x$New_Name_List。

这样数据最后会是这样的:

  EPR_country_code EPR_country           EPR_group_lower_2  EPR_group_lower_3
1              835      Brunei        bumiputera (muslims)  bumiputera (muslims)
2              835      Brunei          bumiputera (other)  bumiputera (other)
3              835      Brunei non bumiputera (indigenous)  non bumiputera
4              835      Brunei                     chinese  chinese

非常感谢

【问题讨论】:

    标签: r recode


    【解决方案1】:

    我们可以使用ifelse 并根据New_Name_List 中的重复值从Old_Name_List 或New_Name_List 中选择值。

    y$EPR_group_lower_3 <- with(x, ifelse(duplicated(New_Name_List) | 
            duplicated(New_Name_List, fromLast = TRUE), Old_Name_List, New_Name_List))
    y
    
    #  EPR_country_code EPR_country           EPR_group_lower_2    EPR_group_lower_3
    #1              835      Brunei        bumiputera (muslims) bumiputera (muslims)
    #2              835      Brunei          bumiputera (other)   bumiputera (other)
    #3              835      Brunei non bumiputera (indigenous)       non bumiputera
    #4              835      Brunei                     chinese              chinese
    

    或者找到值重复的索引并仅替换那些。

    y$EPR_group_lower_3 <- x$New_Name_List
    inds <- with(x, duplicated(New_Name_List) | duplicated(New_Name_List, fromLast = TRUE))
    y$EPR_group_lower_3[inds] <- x$Old_Name_List[inds]
    

    数据

    x <- structure(list(New_Name_List = c("bumiputera", "bumiputera", 
    "non bumiputera", "chinese"), Old_Name_List = c("bumiputera (muslims)", 
    "bumiputera (other)", "non bumiputera (indigenous)", "chinese"
    )), class = "data.frame", row.names = c(NA, -4L))
    
    y <- structure(list(EPR_country_code = c(835L, 835L, 835L, 835L), 
    EPR_country = c("Brunei", "Brunei", "Brunei", "Brunei"), 
    EPR_group_lower_2 = c("bumiputera (muslims)", "bumiputera (other)", 
    "non bumiputera (indigenous)", "chinese")), class = "data.frame", 
    row.names = c(NA, -4L))
    

    【讨论】:

      猜你喜欢
      • 2023-01-19
      • 2022-11-28
      • 2013-06-22
      • 2022-01-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-04-10
      相关资源
      最近更新 更多