【发布时间】:2020-09-11 16:37:12
【问题描述】:
我有一个由 2 列组成的数据集,一列用于用户,一列用于文本:
`User` `Text`
49 there is a cat under the table
21 the sun is hot
431 could you please close the window?
65 there is a cat under the table
21 the sun is hot
53 there is a cat under the table
我的预期输出是:
Text Freq
there is a cat under the table 3
the sun is hot 2
could you please close the window? 1
我的做法是用fuzz.partial_ratio判断所有句子的匹配度(相似度),然后用groupby计算频率。
我使用的是 fuzz.partial_ratio 所以如果完全匹配,它将返回 1(100):
check_match =df.apply(lambda row: ((fuzz.partial_ratio(row['Text'], row['Text'])) >= value), axis=1)
其中 value 是阈值。这是为了确定匹配/相似性
【问题讨论】:
-
这是熊猫数据框吗?
-
是的,它是一个熊猫数据框
-
该方法的代码在哪里?有什么不好的? Stack Overflow 不是编码服务;你必须做出诚实的尝试,然后然后就你的算法或技术提出一个具体的问题。
-
你说你打算使用
fuzz.partial_ratio但在你的例子中你有完全匹配的值 -
@LucaDiMauro 我认为您还应该编辑示例,
fuzzywuzzy在示例中尚不合理,它是value_counts()(因为示例中的记录完全匹配) ,编辑示例以明确何时以及为什么使用fuzz.ratio以及相应的预期输出。
标签: python pandas fuzzywuzzy