【问题标题】:How to count how many sentences are similar?如何计算有多少句子是相似的?
【发布时间】:2020-09-11 16:37:12
【问题描述】:

我有一个由 2 列组成的数据集,一列用于用户,一列用于文本:

`User`        `Text`
49        there is a cat under the table
21        the sun is hot
431       could you please close the window?
65        there is a cat under the table
21        the sun is hot
53        there is a cat under the table

我的预期输出是:

Text                                   Freq         
there is a cat under the table          3
the sun is hot                          2
could you please close the window?      1

我的做法是用fuzz.partial_ratio判断所有句子的匹配度(相似度),然后用groupby计算频率。

我使用的是 fuzz.partial_ratio 所以如果完全匹配,它将返回 1(100):

check_match =df.apply(lambda row: ((fuzz.partial_ratio(row['Text'], row['Text'])) >= value), axis=1)

其中 value 是阈值。这是为了确定匹配/相似性

【问题讨论】:

  • 这是熊猫数据框吗?
  • 是的,它是一个熊猫数据框
  • 该方法的代码在哪里?有什么不好的? Stack Overflow 不是编码服务;你必须做出诚实的尝试,然后然后就你的算法或技术提出一个具体的问题。
  • 你说你打算使用 fuzz.partial_ratio 但在你的例子中你有完全匹配的值
  • @LucaDiMauro 我认为您还应该编辑示例,fuzzywuzzy 在示例中尚不合理,它是value_counts()(因为示例中的记录完全匹配) ,编辑示例以明确何时以及为什么使用 fuzz.ratio 以及相应的预期输出。

标签: python pandas fuzzywuzzy


【解决方案1】:

你可以使用value_counts()

df['Text'].value_counts()

【讨论】:

    【解决方案2】:

    试试这个:

    df = df.groupby('Text').count()
    

    【讨论】:

      【解决方案3】:

      以下应该有效:

      from collections import Counter
      
      l=dict(Counter(df.Text))
      new_df=pd.DataFrame({'Text':list(d.keys()),'Freq': list(d.values())})
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2016-11-22
        • 2017-05-28
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-03-30
        相关资源
        最近更新 更多