【问题标题】:Identifying specific differences between two data sets in R识别 R 中两个数据集之间的特定差异
【发布时间】:2014-12-11 18:09:19
【问题描述】:

我想比较两个数据集并确定它们之间存在差异的具体实例(即哪些变量不同)。

虽然我发现了如何识别两个数据集之间的哪些记录不相同(使用此处详述的函数:http://www.cookbook-r.com/Manipulating_data/Comparing_data_frames/),但我不确定如何标记哪些变量是不同的。

例如

数据集A:

id      name        dob       vaccinedate  vaccinename  dose
100000  John Doe    1/1/2000  5/20/2012    MMR          4
100001  Jane Doe    7/3/2011  3/14/2013    VARICELLA    1

数据集 B:

id      name        dob       vaccinedate  vaccinename  dose
100000  John Doe    1/1/2000  5/20/2012    MMR          3
100001  Jane Doee   7/3/2011  3/24/2013    VARICELLA    1
100002  John Smith  2/5/2010  7/13/2013    HEPB         3

我想确定哪些记录不同,哪些特定变量存在差异。例如,John Doe 记录在 dose 中有 1 个差异,而 Jane Doe 记录在 name 和 vaccinedate 中有 2 个差异。此外,数据集 B 有一个不在数据集 A 中的附加记录,我也想识别这些实例。

最终,目标是找出错误“类型”的频率,例如有多少记录在疫苗日期、疫苗名称、剂量等方面存在差异。

谢谢!

【问题讨论】:

标签: r


【解决方案1】:

这应该可以帮助您入门,但可能还有更优雅的解决方案。

首先,建立df1和df2,以便其他人可以快速复制:

df1 <- structure(list(id = 100000:100001, name = structure(c(2L, 1L), .Label = c("Jane Doe","John Doe"), class = "factor"), dob = structure(1:2, .Label = c("1/1/2000", "7/3/2011"), class = "factor"), vaccinedate = structure(c(2L, 1L), .Label = c("3/14/2013", "5/20/2012"), class = "factor"), vaccinename = structure(1:2, .Label = c("MMR", "VARICELLA"), class = "factor"), dose = c(4L, 1L)), .Names = c("id", "name", "dob", "vaccinedate", "vaccinename", "dose"), class = "data.frame", row.names = c(NA, -2L))

df2 <- structure(list(id = 100000:100002, name = structure(c(2L, 1L, 3L), .Label = c("Jane Doee", "John Doe", "John Smith"), class = "factor"), dob = structure(c(1L, 3L, 2L), .Label = c("1/1/2000", "2/5/2010", "7/3/2011"), class = "factor"), vaccinedate = structure(c(2L, 1L, 3L), .Label = c("3/24/2013", "5/20/2012", "7/13/2013"), class = "factor"), vaccinename = structure(c(2L, 3L, 1L), .Label = c("HEPB", "MMR", "VARICELLA"), class = "factor"), dose = c(3L, 1L, 3L)), .Names = c("id", "name", "dob", "vaccinedate", "vaccinename", "dose"), class = "data.frame", row.names = c(NA, -3L))

接下来,通过mapply 和setdiff 获取从df1 到df2 的差异。也就是说,第一组中没有第二组的内容:

discrep <- mapply(setdiff, df1, df2)
discrep
# $id
# integer(0)
# 
# $name
# [1] "Jane Doe"
# 
# $dob
# character(0)
# 
# $vaccinedate
# [1] "3/14/2013"
# 
# $vaccinename
# character(0)
# 
# $dose
# [1] 4

我们可以使用sapply:

num.discrep <- sapply(discrep, length)
num.discrep
# id        name         dob vaccinedate vaccinename        dose 
# 0           1           0           1           0           1 

根据您关于获取第二组中不在第一组中的 id 的问题,您可以使用 mapply(setdiff, df2, df1) 反转该过程,或者如果它只是 ids 的练习,则只有您可以使用 setdiff(df2$id, df1$id)。

有关 R 的功能函数(例如,mapply、sapply、lapply 等)的更多信息,请参阅this post。


使用purrr 解决方案进行更新:

map2(df1, df2, setdiff) %>% 
  map_int(length)

【讨论】:

    【解决方案2】:

    一种可能性。首先,找出两个数据集共有的 id。最简单的方法是:

    commonID<-intersect(A$id,B$id)
    

    然后您可以通过以下方式确定 A 中缺少哪些行:

    > B[!B$id %in% commonID,]
    #       id       name      dob vaccinedate vaccinename dose
    # 3 100002 John Smith 2/5/2010   7/13/2013        HEPB    3
    

    接下来,您可以将两个数据集限制为它们共有的 id。

    Acommon<-A[A$id %in% commonID,]
    Bcommon<-B[B$id %in% commonID,]
    

    如果你不能假设 id 的顺序是正确的,那么对它们都进行排序:

    Acommon<-Acommon[order(Acommon$id),]
    Bcommon<-Bcommon[order(Bcommon$id),]
    

    现在您可以看到像这样不同的字段。

    diffs<-Acommon != Bcommon
    diffs
    #      id  name   dob vaccinedate vaccinename  dose
    # 1 FALSE FALSE FALSE       FALSE       FALSE  TRUE
    # 2 FALSE  TRUE FALSE        TRUE       FALSE FALSE
    

    这是一个逻辑矩阵,你可以用它做任何你想做的事情。例如,要查找每列中的错误总数:

    colSums(diffs)
    #         id        name         dob vaccinedate vaccinename        dose 
    #          0           1           0           1           0           1 
    

    查找名称不同的所有id:

    Acommon$id[diffs[,"name"]]
    # [1] 100001
    

    等等。

    【讨论】:

    • 谢谢!我没有在上面的示例数据框中指定,但我的实际数据对于每个 id 都有多个记录。例如,John Doe 可以有 5 种疫苗,每种疫苗可以有多个剂量。在您的第一行代码中,我如何确定两个数据集有哪些共同点,而不仅仅是基于 id?希望这是有道理的。
    • 这个问题没有具体的答案。问题是,如果两行不相同,那么您如何确定它们是否“应该”相同但存在差异,或者它们是否实际上是完全不同的条目。您必须提出一些标准才能做出决定。
    • 确实如此。其中一个数据集是“黄金标准”(来自纸质疫苗接种记录),而另一个是单独以电子方式输入的,因此第一个数据集应该是“正确的”数据集。这有助于澄清事情吗?执行此审核的前一个人在 Excel 中手动查看了差异,发现了超过 1000 个错误。理想情况下,我想避免这种手动工作! :)
    【解决方案3】:

    有一个新的包调用 waldo

    install.packages("waldo")
    library(waldo)
    
    # construct the data frames
    
    
    df1 <- structure(list(id = 100000:100001, name = structure(c(2L, 1L), .Label = c("Jane Doe","John Doe"), class = "factor"), dob = structure(1:2, .Label = c("1/1/2000", "7/3/2011"), class = "factor"), vaccinedate = structure(c(2L, 1L), .Label = c("3/14/2013", "5/20/2012"), class = "factor"), vaccinename = structure(1:2, .Label = c("MMR", "VARICELLA"), class = "factor"), dose = c(4L, 1L)), .Names = c("id", "name", "dob", "vaccinedate", "vaccinename", "dose"), class = "data.frame", row.names = c(NA, -2L))
    
    df2 <- structure(list(id = 100000:100002, name = structure(c(2L, 1L, 3L), .Label = c("Jane Doee", "John Doe", "John Smith"), class = "factor"), dob = structure(c(1L, 3L, 2L), .Label = c("1/1/2000", "2/5/2010", "7/3/2011"), class = "factor"), vaccinedate = structure(c(2L, 1L, 3L), .Label = c("3/24/2013", "5/20/2012", "7/13/2013"), class = "factor"), vaccinename = structure(c(2L, 3L, 1L), .Label = c("HEPB", "MMR", "VARICELLA"), class = "factor"), dose = c(3L, 1L, 3L)), .Names = c("id", "name", "dob", "vaccinedate", "vaccinename", "dose"), class = "data.frame", row.names = c(NA, -3L))
    
    # compare them
    compare(df1,df2)
    

    我们得到:

    `old` is length 2
    `new` is length 3
    
    `names(old)`: "X" "Y"    
    `names(new)`: "X" "Y" "Z"
    
    `attr(old, 'row.names')`: 1 2 3  
    `attr(new, 'row.names')`: 1 2 3 4
    
    `old$X`: 1 2 3  
    `new$X`: 1 2 3 4
    
    `old$Y`: "a" "b" "c"    
    `new$Y`: "A" "b" "c" "d"
    
    `old$Z` is absent
    `new$Z` is a character vector ('k', 'l', 'm', 'n')
    

    【讨论】:

      【解决方案4】:
      library(compareDF)
      
      compare_df(dataframe1, dataframe2, c("columnname"))
      

      【讨论】:

        猜你喜欢
        • 2014-01-04
        • 2018-09-28
        • 2011-01-18
        • 1970-01-01
        • 2011-11-11
        • 1970-01-01
        • 1970-01-01
        • 2015-10-12
        • 1970-01-01
        相关资源
        最近更新 更多