【问题标题】:How can I search a text file for a list of words from user input?如何在文本文件中搜索用户输入的单词列表?
【发布时间】:2014-11-23 12:35:17
【问题描述】:

我正在尝试在 Python 3.4.1 中制作一个简单的单词计数器程序,用户将在其中输入以逗号分隔的单词列表,然后在示例文本文件中分析这些单词的频率。

我目前不知道如何在文本文件中搜索输入的单词列表。

首先我尝试过:

file = input("What file would you like to open? ")
f = open(file, 'r')
search = input("Enter the words you want to search for (separate with commas): ").lower().split(",")
search = [x.strip(' ') for x in search]
count = {}
for word in search:
    count[word] = count.get(word,0)+1
for word in sorted(count):
    print(word, count[word])

这导致:

What file would you like to open? twelve_days_of_fast_food.txt
Enter the words you want to search for (separate with commas): first, rings, the
first 1
rings 1
the 1

如果这有什么可做的,我猜这个方法只给了我输入列表中单词的计数,而不是文本文件中单词输入列表的计数。于是我尝试了:

file = input("What file would you like to open? ")
f = open(file, 'r')
lines = f.readlines()
line = f.readline()
word = line.split()
search = input("Enter the words you want to search for (separate with commas): ").lower().split(",")
search = [x.strip(' ') for x in search]
count = {}
for word in lines:
    if word in search:
        count[word] = count.get(word,0)+1
for word in sorted(count):
    print(word, count[word])

这没有给我任何回报。事情是这样的:

What file would you like to open? twelve_days_of_fast_food.txt
Enter the words you want to search for (separate with commas): first, the, rings
>>> 

我做错了什么?我该如何解决这个问题?

【问题讨论】:

    标签: python word-count word-frequency


    【解决方案1】:

    您首先阅读所有行(进入lines,然后尝试仅读取一行,但文件已经为您提供了所有行。在这种情况下,f.readline() 给您一个空行。从在那里,你的剧本注定要失败;你不能计算空行中的单词。

    您可以改为遍历文件:

    file = input("What file would you like to open? ")
    
    search = input("Enter the words you want to search for (separate with commas): ")
    search = [word.strip() for word in search.lower().split(",")]
    
    # create a dictionary for all search words, setting each count to 0
    count = dict.fromkeys(search, 0)
    
    with open(file, 'r') as f:
        for line in f:
            for word in line.lower().split():
                if word in count:
                    # found a word you wanted to count, so count it
                    count[word] += 1
    

    with 语句使用打开的文件对象作为上下文管理器;这只是意味着它会在完成后再次自动关闭。

    for line in f: 循环遍历输入文件中的每一行;这比使用f.readlines() 一次将所有行读入内存更有效。

    我还稍微清理了您的搜索词剥离,并将count字典设置为一个,所有搜索词都预定义为0;这使得实际计数更容易一些。

    因为您现在有一个包含所有搜索词的字典,所以最好针对该字典进行匹配词的测试。对字典进行测试比对列表进行测试更快(后者是列表中的单词越多,扫描需要的时间越长,而字典测试平均需要恒定的时间,无论字典中的项目数量如何)。

    【讨论】:

    • 在 file_.readlines().split(',') 上使用collections.Counter 怎么样?嗯,不,毕竟仍然需要迭代每一行。但也许 collections.Counter(file_.read()) 会派上用场?
    • @brainovergrow:collections.Counter() 是一个很好的补充,但需要导入,并且还会突破 OP 已经熟悉的技术的界限。
    • @brainovergrow: collections.Counter(f.read().lower().split()) 可以,然后查找其中每个搜索词的计数。但首先过滤搜索词也是一种不错的方法,因为这样会占用更少的内存。
    • 对。很高兴知道顺便说一句。在 Python 3.x 中,“文件”不再是内置的 :)(只是将它留给其他 SO 用户,我很确定你已经知道了。)
    • 进行了编辑并对其进行了测试,效果很好,您的解释为我消除了很多困惑。非常感谢,非常感谢!
    【解决方案2】:

    你可以试试这个;

    import re
    import collections
    
    wanted = ["cat", "dog"]
    matches = re.findall('\w+',open('hamlet.txt').read().lower())
    counts = collections.Counter(matches) # Count each occurance of words
    map(lambda x:(x,counts[x]),wanted) # Will print the counts for wanted words
    

    我在形成答案时引用了this solution。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2018-11-27
      • 1970-01-01
      • 2013-05-27
      • 1970-01-01
      • 2014-02-03
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多