【问题标题】:How to merge multiple variables and have one of the variables be in a fuzzy match如何合并多个变量并使其中一个变量处于模糊匹配中
【发布时间】:2021-05-05 14:34:42
【问题描述】:

我最初是在之前的帖子中帮助我进行模糊匹配

I would like to match two datasets based on arbitrary address fields relative to each other using R

感谢@Ronak Shah、@r2evans 和 @akrun 提供以前的帮助

这很有帮助,我根据这两个数据集得到了我想要的模糊匹配

structure(list(ID = 1:8, Address = c("Canal and Broadway", "55 water street room number 73", 
"Mulberry street", "Front street and Fulton", "62nd street ", 
"wythe street", "vanderbilt avenue", "South Beach avenue")), class = "data.frame", row.names = c(NA, 
-8L))

和

structure(list(ID2 = 1:8, Address = c("Canal & Broadway", "Somewhere around 55 water street", 
"Mulberry street", "Front street and close to Fulton", "south beach avenue", 
"along wythe street on the southwest ", "vanderbilt ave", "62nd street"
)), class = "data.frame", row.names = c(NA, -8L))

运行

fuzzyjoin::stringdist_left_join(df1, df2, by = 'Address', max_dist = 5)

给我

structure(list(ID = 1:8, Address.x = c("Canal and Broadway", 
"55 water street room number 73", "Mulberry street", "Front street and Fulton", 
"62nd street ", "wythe street", "vanderbilt avenue", "South Beach avenue"
), ID2 = c(1L, NA, 3L, NA, 8L, 8L, 7L, 5L), Address.y = c("Canal & Broadway", 
NA, "Mulberry street", NA, "62nd street", "62nd street", "vanderbilt ave", 
"south beach avenue")), row.names = c(NA, -8L), class = "data.frame")

比赛做得很好,我接受。我接下来要做的是匹配 df1_new 和 df2_new

df1

structure(list(ID = 1:8, Address = c("Canal and Broadway", "55 water street room number 73", 
"Mulberry street", "Front street and Fulton", "62nd street ", 
"wythe street", "vanderbilt avenue", "South Beach avenue"), Age = c(32L, 
33L, 37L, 39L, 38L, 50L, 60L, 42L), Name = c("John ", "Adam", 
"Alan", "Greg", "Phil", "Anthony", "Mike", "Mark")), class = "data.frame", row.names = c(NA, 
 -8L))

和 df 2

structure(list(ID2 = 1:8, Address = c("Canal & Broadway", "Somewhere around 55 water street", 
"Mulberry street", "Front street and close to Fulton", "south beach avenue", 
"along wythe street on the southwest ", "vanderbilt ave", "62nd street"
), Age = c(32L, 33L, 37L, 39L, 42L, 50L, 60L, 35L), Name = c("John", 
"Adam", "Ryan", "Greg", "Mark", "Anthony", "Mike", "Phil")), class = "data.frame", row.names = c(NA,-8L))

通常我会跑

df3<-df1 %>% left_join(df2, by=c("Address","Age","Name")

但是,地址变量需要经过模糊匹配,而其他变量可以经过标准。我希望把left_join和fuzzyjoin::stringdist_left_join函数放在一起

ID   Address.x                       D2   Address.y        Age        Name
1   Canal and Broadway                1   Canal & Broadway  32        John
2   55 water street room number 73    
3   Mulberry street                   
4   Front street and Fulton           
5   62nd street                       8 62nd street
6   wythe street                      
7   vanderbilt avenue                 7 vanderbilt ave      60        Mike
8   South Beach avenue                5 south beach avenue  42        Mark

请注意,虽然 62nd street 和 Mulberry street 在模糊匹配中匹配,但它们对应的 Age 和 Name 并不相同。

【问题讨论】:

标签: r join dplyr


【解决方案1】:
fuzzyjoin::stringdist_left_join(df1_new, df2_new ['Address'], by = 'Address', max_dist 
= 5) %>%
mutate(Address.z=Address.y) %>% left_join(df2_new %>% 
mutate(Address.z=Address),by=c("Age","Name", "Address.z"))

这让我得到了我想要的结果。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-01-17
    • 2012-01-15
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多