很难找到一种万能的解决方案。这是因为看起来相似的字符串可能描述了非常不同的事物(例如格拉纳达与格林纳达)。原帖下的cmets值得研究。
参见"Approximate string matching" on Wikipedia(有时也称为“模糊匹配”)。如您所见,有很多方法可以在字符串上定义“相似”。
最基本的工具是R函数adist。它计算所谓的编辑距离。
x <- c("American Indian and Alaska Native" ,
"Asian" ,
"Black of African American" ,
"Black or African American" ,
"Other" ,
"Unknown" ,
"white or Caucasian" ,
"White or Caucasian" ,
"White or Caucasion" )
u <- unique(x)
# compare all strings against each other
d <- adist(u)
# Do not list combinations of similar words twice
d[lower.tri(d)] <- NA
# Say your threshold below which you want to consider strings similar is
# 2 edits:
a <- which(d > 0 & d < 2, arr.ind = TRUE)
a
## row col
## [1,] 3 4
## [2,] 7 8
## [3,] 8 9
pairs <- cbind(u[a[,1]], u[a[,2]])
pairs
## [,1] [,2]
## [1,] "Black of African American" "Black or African American"
## [2,] "white or Caucasian" "White or Caucasian"
## [3,] "White or Caucasian" "White or Caucasion"
但最终,您必须自己策划结果,以避免不公平因素的意外均衡。
您可以通过使用命名向量作为翻译字典来重复执行此操作。例如,通过查看上面的示例,我可以创建以下字典:
dict <- c(
# incorrect spellings correct spellings
# ------------------------- ----------------------------
"Black of African American" = "Black or African American",
"white or Caucasian" = "white or Caucasian" ,
"White or Caucasion" = "White or Caucasian"
)
# The correct levels need to be included, to
dict <- c(dict, setNames(u,u)
然后使用as.character 将您的因子列转换为字符并应用
上面的字典就像我在这里用原始字符向量x:
xcorrected <- dict[x]
# show without names, but the result is also correct if you just use
# xcorrected alone (remove as.character here to see the difference).
as.character(xcorrected)
[1] "American Indian and Alaska Native" "Asian"
[3] "Black or African American" "Black or African American"
[5] "Other" "Unknown"
[7] "white or Caucasian" "White or Caucasian"
[9] "White or Caucasian"