有没有更好的方法从 python 列表中获取“重要单词”？答案

【问题标题】：Is there a better way to get just 'important words' from a list in python?有没有更好的方法从 python 列表中获取“重要单词”？
【发布时间】：2013-08-15 06:39:33
【问题描述】：

我使用reddit praw api编写了一些代码来查找reddit上提交标题中最流行的单词。

import nltk
import praw

picksub = raw_input('\nWhich subreddit do you want to analyze? r/')
many = input('\nHow many of the top words would you like to see? \n\t> ')

print 'Getting the top %d most common words from r/%s:' % (many,picksub)
r = praw.Reddit(user_agent='get the most common words from chosen subreddit')
submissions = r.get_subreddit(picksub).get_top_from_all(limit=200)

hey = []

for x in submissions:
    hey.extend(str(x).split(' '))   

fdist = nltk.FreqDist(hey) # creates a frequency distribution for words in 'hey'
top_words = fdist.keys()

common_words = ['its','am', 'ago','took', 'got', 'will', 'been', 'get', 'such','your','don\'t', 'if', 'why', 'do', 'does', 'or', 'any', 'but', 'they', 'all', 'now','than','into','can', 'i\'m','not','so','just', 'out','about','have','when', 'would' ,'where', 'what', 'who' 'I\'m','says' 'not', '', 'over', '_', '-','after', 'an','for', 'who', 'by', 'from', 'it', 'how', 'you', 'about' 'for', 'on', 'as', 'be', 'has', 'that', 'was', 'there', 'with','what', 'we', '::', 'to', 'the', 'of', ':', '...', 'a', 'at', 'is', 'my', 'in' , 'i', 'this', 'and', 'are', 'he', 'she', 'is', 'his', 'hers']
already = []
counter = 0
number = 1

print '-----------------------'
for word in top_words:  
    if word.lower() not in common_words and word.lower() not in already:
        print str(number) + ". '" + word + "'"
        counter +=1
    number +=1
    already.append(word.lower())
if counter == many:
    break
print '-----------------------\n'

所以输入 subreddit 'python' 并获得 10 个帖子返回：

'Python'
'PyPy'
'代码'
'使用'
'136'
'181'
'd...'
'IPython'
'133'
10. '158'

我怎样才能让这个脚本不返回数字，以及像“d...”这样的错误词？前 4 个结果是可以接受的，但我想用有意义的词代替其余的。列出 common_words 是不合理的，并且不会过滤这些错误。我对编写代码比较陌生，感谢您的帮助。

【问题讨论】：

标签： python api nltk reddit

【解决方案1】：

我不同意。制作常用词列表是正确的，没有更简单的方法可以过滤掉，for，i，am等。但是，使用common_words列表过滤掉不是词的结果是不合理的，因为那时你必须包括你不想要的所有可能的非单词。应该以不同的方式过滤掉非单词。

一些建议：
1) common_words 应该是set()，因为你的列表很长，这应该会加快速度。 in 对集合的操作是 O(1)，而对于列表是 O(n)。

2) 摆脱所有数字字符串是微不足道的。一种方法是：

all([w.isdigit() for w in word])

如果返回 True，那么这个词只是一系列数字。

3) 摆脱 d... 有点棘手。这取决于您如何定义非单词。这个：

tf = [ c.isalpha() for c in word ]

返回 True/False 值列表（如果 char 不是字母，则返回 False）。然后，您可以计算如下值：

t = tf.count(True)
f = tf.count(False)

然后，您可以将非单词定义为其中包含的非字母字符多于字母的单词，以及根本包含任何非字母字符的单词，等等。例如：

def check_wordiness(word):
    # This returns true only if a word is all letters
    return all([ c.isalpha() for c in word ])

4) 在for word in top_words: 块中，你确定你没有混淆计数器和数字吗？此外，计数器和数字几乎是多余的，您可以将最后一位重写为：

for word in top_words:
    # Since you are calling .lower() so much, 
    # you probably want to define it up here
    w = word.lower() 
    if w not in common_words and w not in already:
        # String formatting is preferred over +'s
        print "%i. '%s'" % (number, word)
        number +=1
    # This could go under the if statement. You only want to add
    # words that could be added again.  Why add words that are being
    # filtered out anyways?
    already.append(w)

    # this wasn't indented correctly before
    if number == many:
        break

希望对您有所帮助。

【讨论】：

你可以直接说word.isdigit()，而不是all([w.isdigit() for w in word])。 .isalpha() 也一样。
好点。我决定扩展逻辑，因为 OP 说他是编码新手，列表理解并没有比这更清楚了。

10. '158'