【问题标题】:How to run through list of keyword vectors and fuzzy match them to a different file (R)如何遍历关键字向量列表并将它们模糊匹配到不同的文件(R)
【发布时间】:2018-10-25 15:33:46
【问题描述】:

我有两个文件,一个包含关键字(大约 2,000 行),另一个包含文本(大约 770,000 行)。关键字文件如下所示:

Event Name            Keyword
All-day tabby fest    tabby, all-day
All-day tabby fest    tabby, fest
Maine Coon Grooming   maine coon, groom    
Maine Coon Grooming   coon, groom

keywordFile <- tibble(EventName = c("All-day tabby fest", "All-day tabby fest", "Maine Coon Grooming","Maine Coon Grooming"), Keyword = c("tabby, all-day", "tabby, fest", "maine coon, groom", "coon, groom")

文本文件如下所示:

Description
Bring your tabby to the fest on Tuesday
All cats are welcome to the fest on Tuesday
Mainecoon grooming will happen at noon Wednesday
Maine coons will be pampered at noon on Wednesday

text <- tibble(Description = c("Bring your tabby to the fest on Tuesday","All cats are welcome to the fest on Tuesday","Mainecoon grooming will happen at noon Wednesday","Maine coons will be pampered at noon on Wednesday")

我想要的是遍历文本文件并查找模糊匹配(必须包括“关键字”列中的每个单词)并返回一个显示 TRUE 或 False 的新列。如果这是真的,那么我想要第三列显示事件名称。所以看起来像:

Description                                          Match?   Event Name
Bring your tabby to the fest on Tuesday              TRUE     All-day tabby fest
All cats are welcome to the fest on Tuesday          FALSE
Mainecoon grooming will happen at noon Wednesday     TRUE     Maine Coon Grooming
Maine coons will be pampered at noon on Wednesday    FALSE

感谢 Molx (How can I check if multiple strings exist in another string?),我能够使用这样的东西成功地进行模糊匹配(在将所有内容转换为小写之后):

str <- c("tabby", "all-day")
myStr <- "Bring your tabby to the fest on Tuesday"
all(sapply(str, grepl, myStr))

但是,当我尝试模糊匹配整个文件时,我遇到了困难。我尝试过这样的事情:

for (i in seq_along(text$Description)){
  for (j in seq_along(keywordFile$EventName)) {
    # below I am creating the TRUE/FALSE column
    text$TF[i] <- all(sapply(keywordFile$Keyword[j], grepl, 
                                                     text$Description[i]))
    if (isTRUE(text$TF))
      # below I am creating the EventName column
      text$EventName <- keywordFile$EventName
    }
}

我认为将正确的内容转换为向量和字符串时没有问题。我的keywordFile$Keyword 列是一堆字符串向量,而我的text$Description 列是一个字符串。但是我正在努力解决如何正确遍历这两个文件。我得到的错误是

Error in ... replacement has 13 rows, data has 1

以前有人做过这样的事吗?

【问题讨论】:

  • 我相信我有 (stackoverflow.com/questions/51733851/…),但是如果您以可重现的格式发布数据,它会更容易提供帮助。查看stackoverflow.com/questions/5963269/…
  • 我检查了那里的链接,但是当我尝试实现它时出现错误(模式的长度> 1,只会使用第一个元素)。我已经编辑了我的问题以包含一些代码以及我在运行 for 循环代码时遇到的错误。我希望这有帮助。真的坚持这一点。

标签: r loops matching sapply grepl


【解决方案1】:

我不完全确定我明白你的问题,因为我不会调用 grepl() 模糊匹配。如果关键字在较长的单词中,它会更愿意捕获它。所以“猫”和“灾难”将是一个匹配事件,认为这些词非常不同。

我选择写一个答案是你可以控制仍然构成匹配的字符串之间的距离:

加载库:

library(tibble)
library(dplyr)
library(fuzzyjoin)
library(tidytext)
library(tidyr)

制作字典和数据对象:

dict <- tibble(Event_Name = c(
  "All-day tabby fest",
  "All-day tabby fest",
  "Maine Coon Grooming",
  "Maine Coon Grooming"
), Keyword = c(
  "tabby, all-day",
  "tabby, fest",
  "maine coon, groom",
  "coon, groom"
)) %>% 
  mutate(Keyword = strsplit(Keyword, ", ")) %>% 
  unnest(Keyword)

string <- tibble(id = 1:4, Description = c(
  "Bring your tabby to the fest on Tuesday",
  "All cats are welcome to the fest on Tuesday",
  "Mainecoon grooming will happen at noon Wednesday",
  "Maine coons will be pampered at noon on Wednesday"
))

将字典应用于数据:

string_annotated <- string %>% 
  unnest_tokens(output = "word", input = Description) %>%
  stringdist_left_join(y = dict, by = c("word" = "Keyword"), max_dist = 1) %>% 
  mutate(match = !is.na(Keyword))

> string_annotated
# A tibble: 34 x 5
      id word    Event_Name         Keyword match
   <int> <chr>   <chr>              <chr>   <lgl>
 1     1 bring   NA                 NA      FALSE
 2     1 your    NA                 NA      FALSE
 3     1 tabby   All-day tabby fest tabby   TRUE 
 4     1 tabby   All-day tabby fest tabby   TRUE 
 5     1 to      NA                 NA      FALSE
 6     1 the     NA                 NA      FALSE
 7     1 fest    All-day tabby fest fest    TRUE 
 8     1 on      NA                 NA      FALSE
 9     1 tuesday NA                 NA      FALSE
10     2 all     NA                 NA      FALSE
# ... with 24 more rows

max_dist 控制仍然构成匹配的内容。在这种情况下,1 或更小的字符串之间的距离可以找到所有文本的匹配项,但我也尝试使用不匹配的字符串。

如果您想将这种长格式恢复为原始格式:

string_annotated_col <- string_annotated %>% 
  group_by(id) %>% 
  summarise(Description = paste(word, collapse = " "),
            match = sum(match),
            keywords = toString(unique(na.omit(Keyword))),
            Event_Name = toString(unique(na.omit(Event_Name))))

> string_annotated_col
# A tibble: 4 x 5
     id Description                                       match keywords         Event_Name         
  <int> <chr>                                             <int> <chr>            <chr>              
1     1 bring your tabby tabby to the fest on tuesday         3 tabby, fest      All-day tabby fest 
2     2 all cats are welcome to the fest on tuesday           1 fest             All-day tabby fest 
3     3 mainecoon grooming will happen at noon wednesday      2 maine coon, coon Maine Coon Grooming
4     4 maine coons will be pampered at noon on wednesday     2 coon             Maine Coon Grooming

如果部分答案对您没有意义,请随时提出问题。其中一些在here 中进行了解释。除了模糊匹配部分。

【讨论】:

  • 我为迟到的回复道歉;我一直在尝试用我的数据来解决这个问题,它似乎工作得很好。还必须阅读一下这里到底发生了什么,所以感谢你把所有的东西都写出来。我真的很感激这一点。非常感谢。
  • 没问题。这是一个有趣的挑战,我在这个过程中发现了 fuzzyjoin 包,这与我自己的工作非常相关
【解决方案2】:

可以在 R 中使用 agrep()grepl() 函数进行近似匹配。它适用于选项fixed=False。这些函数不需要任何额外的库。

【讨论】:

    猜你喜欢
    • 2023-03-19
    • 1970-01-01
    • 2022-11-20
    • 1970-01-01
    • 2016-07-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-03-22
    相关资源
    最近更新 更多