【问题标题】:Recycling error while using stringdist and data.table in R在 R 中使用 stringdist 和 data.table 时出现回收错误
【发布时间】:2019-06-12 13:11:16
【问题描述】:

我正在尝试对包含作者姓名的数据表执行近似字符串匹配,该数据表基于“名字”字典。我还设置了一个高于 0.9 的高阈值,以提高匹配质量。

但是,我收到以下错误消息:

Warning message:
In [`<-.data.table`(x, j = name, value = value) :
Supplied 6 items to be assigned to 17789 items of column 'Gender_Dict' (recycled leaving remainder of 5 items).

即使我使用 signif(similarity_score,4) 将相似度匹配向下舍入到 4 位,也会出现此错误。

有关输入数据和方法的更多信息:

  1. author_corrected_df 是一个 data.table,包含以下列:“Author”和“Author_Corrected”。 Author_Corrected 是相应作者的字母表示形式(例如:如果 Author = Jack123,则 Author_Corrected = Jack)。
  2. Author_Corrected 列可以有正确名字的变体,例如:Jackk 而不是 Jack,我想在名为 Gender_Dict 的 author_corrected_df 中填充相应的性别。
  3. 另一个名为 first_names_dict 的 data.table 包含“姓名”(即名字)和性别(0 表示女性,1 表示男性,2 表示关系)。
  4. 我想从每行的“Author_Corrected”中找到与 first_names_dict 中的“姓名”最相关的匹配项,并填充相应的性别(0、1、2 之一)。
  5. 为了使字符串匹配更加严格,我使用了 0.9720 的阈值,否则在后面的代码中(未在下面显示),不匹配的值则表示为 NA。
  6. 可以从以下链接访问 first_names_dict 和 author_corrected_df: https://wetransfer.com/downloads/6efe42597519495fcd2c52264c40940a20190612130618/0cc87541a9605df0fcc15297c4b18b7d20190612130619/6498a7
for (ijk in 1:nrow(author_corrected_df)){
  max_sim1 <- max(stringsim(author_corrected_df$Author_Corrected[ijk], first_names_dict$name, method = "jw", p = 0.1, nthread = getOption("sd_num_thread")), na.rm = TRUE)
  if (signif(max_sim1,4) >= 0.9720){
    row_idx1 <- which.max(stringsim(author_corrected_df$Author_Corrected[ijk], first_names_dict$name, method = "jw", p = 0.1, nthread = getOption("sd_num_thread")))
    author_corrected_df$Gender_Dict[ijk] <- first_names_dict$gender[row_idx1]
  } else {
    next
  }
}

执行时我收到以下错误消息:

Warning message:
In `[<-.data.table`(x, j = name, value = value) :
  Supplied 6 items to be assigned to 17789 items of column 'Gender_Dict' (recycled leaving remainder of 5 items).

在了解错误所在以及是否有更快的方法来执行此类匹配(尽管后者是第二优先级)方面,我们将不胜感激。

提前致谢。

【问题讨论】:

  • 您好,我建议您在定义这些变量后添加print(max_sim1) 和print(row_idx1)。
  • 嗨,cbo,我尝试将打印语句添加到变量中,但无法弄清楚这有什么用处。示例输出如下所示: 1 114654 1 114654 0.95 0.9333333 0.9333333 0.925 0.9142857 0.93 0.8933333 但我仍然得到与上述相同的错误。
  • 这确认了6 items to be assigned to 17789,您需要一对一的映射。通过在没有循环的情况下运行代码来检查您是否有多个最大值(例如,使用 ijk max_sim1、max(stringsim(author_corrected_df$...、which.max(stringsim(author_corrected_df$...、author_corrected_df$Gender_Dict[ijk]、first_names_dict$gender[row_idx1]的输出。
  • 您可以检查row_idx1 并打印可能出现问题的位置,并从所有索引中仅取一个值(例如通过统计数据)。

标签: r data.table stringdist


【解决方案1】:

继之前的 cmets 之后,我在这里选择您选择中最常见的性别:

for (ijk in 1:nrow(author_corrected_df)){
        max_sim1 <- max(stringsim(author_corrected_df$Author_Corrected[ijk], first_names_dict$name, method = "jw", p = 0.1, nthread = getOption("sd_num_thread")), na.rm = TRUE)
        if (signif(max_sim1,4) >= 0.9720){
                row_idx1 <- which.max(stringsim(author_corrected_df$Author_Corrected[ijk], first_names_dict$name, method = "jw", p = 0.1, nthread = getOption("sd_num_thread")))

                # Analysis of factor gender
                gender <- as.character( first_names_dict$gender[row_idx1] )

                # I take the (first) gender most present in selection 
                df_count <- as.data.frame( table(gender) )
                ref <- as.character ( df_count$test[which.max(df_count$Freq)] )
                value <- unique ( test[which(test == ref)] )

                # Affecting single character value to data frame
                author_corrected_df$Gender_Dict[ijk] <- value
        }
}

希望这会有所帮助:)

【讨论】:

  • 同意“问题”中同一作者的许多匹配点。因此,我修改了我的代码,现在将stringsim结果存储为data.table,有2列:首先作为运行索引/计数器,第二作为特定作者的相似度得分,然后我对这个data.table进行排序相似度降序排列,只保留第一行。我的假设是,即使在多个匹配记录的每种情况下,这也只会保留一个实例。但仍然得到同样的错误。这次会检查有什么不工作的!感谢 cbo 的解决方案,我也试试!
  • 好的,这与我当时对 df_count 所做的类似。您还可以使用max 以外的其他统计信息替换`ref wich.max。干杯!
猜你喜欢
  • 1970-01-01
  • 2022-01-08
  • 1970-01-01
  • 2020-07-16
  • 2016-09-02
  • 2017-10-18
  • 2023-04-03
  • 1970-01-01
  • 2020-08-30
相关资源
最近更新 更多