【问题标题】:How to merge two dataframes with same column name but may have same data in variables in R?如何合并具有相同列名但在 R 中的变量中可能具有相同数据的两个数据框?
【发布时间】:2020-07-20 22:18:36
【问题描述】:

我想问一下如何合并这两个数据框?

df1:

Name   Type   Price
A       1      NA
B       2      2.5
C       3      2.0

df2:

Name   Type   Price
A       1      1.5
D       2      2.5
E       3      2.0

正如您从两个 df 中看到的那样,它们具有相同的列名和一行具有相同值的“名称”,即 A,但 df1 没有价格,而 df2 有。我想实现这个输出,如果“名称”中的值相同,它们就会合并

Name   Type   Price
A       1      1.5
B       2      2.5
C       3      2.0
D       2      2.5
E       3      2.0

【问题讨论】:

    标签: r


    【解决方案1】:

    我们可以通过Name 对df1 和df2 执行full_join 并在Type 和Price 上使用coalesce 从这些列中获取第一个非NA 值。

    library(dplyr)
    
    full_join(df1, df2, by = 'Name') %>%
       mutate(Type = coalesce(Type.x, Type.y), 
              Price = coalesce(Price.x, Price.y)) %>%
       select(names(df1))
    
    #  Name Type Price
    #1    A    1   1.5
    #2    B    2   2.5
    #3    C    3   2.0
    #4    D    2   2.5
    #5    E    3   2.0
    

    在基础 R 中类似:

    transform(merge(df1, df2, by = 'Name', all = TRUE), 
               Price = ifelse(is.na(Price.x), Price.y, Price.x), 
               Type = ifelse(is.na(Type.x), Type.y, Type.x))[names(df1)]
    

    数据

    df1 <- structure(list(Name = structure(1:3, .Label = c("A", "B", "C"
    ), class = "factor"), Type = 1:3, Price = c(NA, 2.5, 2)), 
    class = "data.frame", row.names = c(NA, -3L))
    
    df2 <- structure(list(Name = structure(1:3, .Label = c("A", "D", "E"
    ), class = "factor"), Type = 1:3, Price = c(1.5, 2.5, 2)), 
    class = "data.frame", row.names = c(NA, -3L))
    

    【讨论】:

      【解决方案2】:

      您似乎想将数据框重新绑定在一起,然后删除价格为 NA 值的行,并按名称排序。

      library(data.table)
      
      setDT(rbind(df1, df2))[!is.na(Price)][order(Name)]
      #    Name Type Price
      # 1:    A    1   1.5
      # 2:    B    2   2.5
      # 3:    C    3   2.0
      # 4:    D    2   2.5
      # 5:    E    3   2.0
      

      【讨论】:

        【解决方案3】:

        这是使用 merge + ocmplete.cases 的基本 R 解决方案

        dfout <- subset(u <- merge(df1,df2,all= TRUE),complete.cases(u))
        

        产生

        > dfout
          Name Type Price
        1    A    1   1.5
        3    B    2   2.5
        4    C    3   2.0
        5    D    2   2.5
        6    E    3   2.0
        

        数据

        df1 <- structure(list(Name = structure(1:3, .Label = c("A", "B", "C"
        ), class = "factor"), Type = 1:3, Price = c(NA, 2.5, 2)), 
        class = "data.frame", row.names = c(NA, -3L))
        
        df2 <- structure(list(Name = structure(1:3, .Label = c("A", "D", "E"
        ), class = "factor"), Type = 1:3, Price = c(1.5, 2.5, 2)), 
        class = "data.frame", row.names = c(NA, -3L))
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2016-08-12
          • 1970-01-01
          • 2021-03-20
          • 2017-06-24
          • 1970-01-01
          • 2021-03-11
          • 2013-12-08
          • 1970-01-01
          相关资源
          最近更新 更多