【问题标题】:keeping the best string matched by fuzzy matching in R在R中保持通过模糊匹配匹配的最佳字符串
【发布时间】:2020-10-12 11:52:05
【问题描述】:

我在 R 中有两个数据框。一个是我想要匹配的短语的数据框以及它们在另一列 (df.word) 中的同义词,另一个是我想要匹配的字符串的数据框想要与代码(df.string)一起匹配。字符串很复杂,但为了方便起见,我们有:

df.word <- data.frame(label = c('warm wet', 'warm dry', 'cold wet'),
                 synonym = c('hot and drizzling\nsunny and raining','sunny and clear sky\ndry sunny day', 'cold winds and raining\nsnowing'))

df.string <- data.frame(day = c(1,2,3,4),
                       weather = c('there would be some drizzling at dawn but we will have a hot day', 'today there are cold winds and a bit of raining or snowing at night', 'a sunny and clear sky is what we have today', 'a warm dry day'))

我想创建 df.string$extract,我想在其中为字符串提供 最佳匹配。

这样的专栏

df$extract <- c('warm wet', 'cold wet', 'warm dry', 'warm dry')

提前感谢任何人的帮助。

【问题讨论】:

  • 实际上最好的匹配是最长的匹配......在字符串中检测到的单词数量最多的匹配。

标签: r string string-matching fuzzyjoin


【解决方案1】:

你的问题有几点我不太明白;但是,我正在为您的问题提出解决方案。检查它是否适合您。

我假设您想为天气文本找到最匹配的标签。如果是这样,您可以通过以下方式使用来自library(stringdist) 的stringsim 函数。

第一注意:如果你清理数据中的\n,结果会更准确。所以,我在这个例子中清理它们,但如果你愿意,你可以保留它们。

第二个注意:你可以根据不同的方法来改变相似距离。这里我使用了余弦相似度,这是一个比较好的起点。如果您想查看替代方法,请参阅函数的参考:

?stringsim

干净的数据如下:

df.word <- data.frame(
    label = c("warm wet", "warm dry", "cold wet"),
    synonym = c(
        "hot and drizzling sunny and raining",
        "sunny and clear sky dry sunny day", 
        "cold winds and raining snowing"
    )
)

df.string <- data.frame(
    day = c(1, 2, 3, 4),
    weather = c(
        "there would be some drizzling at dawn but we will have a hot day",
        "today there are cold winds and a bit of raining or snowing at night", 
        "a sunny and clear sky is what we have today", 
        "a warm dry day"
    )
)

安装库并加载它

install.packages('stringdist')
library(stringdist)

创建一个n x m 矩阵,其中包含每个是否文本与每个同义词的相似度得分。行显示每个文本和列是否代表每个同义词组。

match.scores <- sapply(          ## Create a nested loop with sapply
    seq_along(df.word$synonym),  ## Loop for each synonym as 'i'
    function(i) {
        sapply(
            seq_along(df.string$weather), ## Loop for each weather as 'j'
            function(j) {
                stringsim(df.word$synonym[i], df.string$weather[j], ## Check similarity 
                    method = "cosine", ## Method cosine  
                    q = 2 ## Size of the q -gram: 2 
                )
            }
        )
    }
)

r$> match.scores
          [,1]      [,2]       [,3]
[1,] 0.3657341 0.1919924 0.24629819
[2,] 0.6067799 0.2548236 0.73552828
[3,] 0.3333974 0.6300619 0.21791793
[4,] 0.1460593 0.4485426 0.03688556

获取每个文本行的最佳匹配,找到匹配分数最高的标签,并将这些标签添加到数据框中。

ranked.match <- apply(match.scores, 1, which.max)
df.string$extract <- df.word$label[ranked.match]

df.string

r$> df.string
  day                                                             weather  extract
1   1    there would be some drizzling at dawn but we will have a hot day warm wet
2   2 today there are cold winds and a bit of raining or snowing at night cold wet
3   3                         a sunny and clear sky is what we have today warm dry
4   4                                                      a warm dry day warm dry

【讨论】:

  • 感谢您的完整解释和回答。我尝试了不同“q”的代码......最好的是q = 5,但仍然不到一半是正确的......我希望标签/同义词的所有单词都存在,但这不会发生在任何方法或 q(我的意思是,例如,我什至得到一个“下雪”天气的“温暖干燥”标签)。
  • 这会很好.. [链接] (stackoverflow.com/questions/62411327/…) 但问题是在这种情况下我只想要最佳答案而不是所有匹配。
  • 我明白了。您能否提交一个最终工作的示例。因此,我可以更好地了解您到底想要什么。您希望在提取的列中看到什么?如果你不介意举个例子。
  • 我不介意,只是文本是波斯语的……它实际上是传统医学文本中某些草药的气质文本……例如,我有一种很热的草药在二级和三级湿,我需要的是“热湿”......或者像一种“平衡微热和一级干燥”的药物,我需要标签“平衡热干”.. . 但是我使用上面的代码,我有一个标签 balanced hot dry 用于一种药物,它只说一个简短的句子“它是干燥的”。我不知道balanced和hot是从哪里来的。
猜你喜欢
  • 2015-03-26
  • 2015-11-10
  • 1970-01-01
  • 2016-08-17
  • 2012-02-14
  • 2014-11-02
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多