【问题标题】:Using nltk library to extract keywords使用 nltk 库提取关键字
【发布时间】:2011-06-08 10:27:24
【问题描述】:

我正在开发一个应用程序,该应用程序需要我从对话流中提取关键字(并最终生成这些词的标签云)。我正在考虑以下步骤:

  1. 标记每个原始对话(输出存储为字符串列表)
  2. 删除停用词
  3. 使用词干分析器(Porter 词干提取算法)

到目前为止,nltk 提供了我需要的所有工具。在此之后,但是我需要以某种方式对这些单词进行“排名”并提出最重要的单词。谁能建议我 nltk 中的哪些工具可以用于此?

谢谢 尼特

【问题讨论】:

  • 一种很有前途的术语排名方法,特别是用于生成词云,是简约的语言模型。在github.com/larsmans/weighwords (WIP) 上查看我的实现

标签: python tags cloud nltk


【解决方案1】:

我想这取决于您对“重要”的定义。 如果您在谈论频率,那么您可以使用单词(或词干)作为键来构建字典,然后将其计为值。之后,您可以根据它们的计数对字典中的键进行排序。

类似的东西(未测试):

from collections import defaultdict

#Collect word statistics
counts = defaultdict(int) 
for sent in stemmed_sentences:
   for stem in sent:
      counts[stem] += 1

#This block deletes all words with count <3
#They are not relevant and sorting will be way faster
pairs = [(x,y) for x,y in counts.items() if y >= 3]

#Sort (stem,count) pairs based on count 
sorted_stems = sorted(pairs, key = lambda x: x[1])

【讨论】:

  • ... 你可以尝试用 idf 惩罚所有太常见的词,尽管用户研究表明 tf cloud 比 tf-idf 更受欢迎。 +1。
  • 还可以查看信息增益指标和显着性测试。 nltk.metrics 在这方面提供了一些很好的功能。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2020-12-02
  • 2011-08-06
  • 1970-01-01
  • 2022-07-31
  • 2015-07-28
  • 1970-01-01
  • 2013-12-15
相关资源
最近更新 更多