新热点
编辑:根据我的回答中的 cmets,OP 正在寻求为文件中的所有单词(空格分隔)构建一个包含 {word:wordcount} 的 dict。
有一个非常棒的方法可以做到这一点,但它并没有真正教给你任何东西,所以我会先向你展示缓慢的方法,然后再包括最佳解决方案。
wordcountdict = dict()
r = input("filename: ")
with open(r, 'r') as infile:
for line in infile:
for word in infile.split(): # split on whitespace
try:
wordcountdict[word.lower()] += 1
# try adding one to the word in the counter
except KeyError:
wordcountdict[word.lower()] = 1
# If the word isn't in the dict already, set it to 1
现在您可能想要过滤掉一些常用词("at"、"I"、"then" 等),在这种情况下,您可以建立它们的黑名单(例如 blacklist = ['at', 'i', 'then'])并执行 if word.lower() in blacklist: continue在for word in infile.split() 内部和try/except 块之前。这将测试该单词是否在黑名单中,如果是则跳过该执行的其余部分。
现在我向你保证了一个很好的方法来做到这一点,那就是collections.Counter。它是专门为计算列表中的元素而创建的字典。有更快的方法来计算单词,但在 Python (imo) 中没有更干净的方法。你在this question查看时间安排
from collections import Counter
wordcountdict = Counter()
r = input("filename: ")
with open(r, 'r') as infile:
for line in infile:
wordcountdict += Counter( map(str.lower,line.split()) )
如果您从未使用过来自 collections 或 map 函数的导入,那么这将是非常神秘的,这就是我没有把它放在首位的原因! :)。
基本上:collections.Counter 将一个可迭代对象作为参数,并计算可迭代对象中的所有元素(因此 `Counter([1,1,2,3,4,4,4]) == {1:2, 2:1、3:1、4:3})。您可以添加它们,它会在它们唯一的地方创建新的键,并在它们不唯一的地方添加值。
map(callable, iterable) 运行 callable 并带有可迭代的每个元素的参数,并返回一个本身可迭代的 map 对象(在 Python2 中是 list)(因此 map(str.lower, ["ThIS", "Has", "UppEr", "aNd", "LOWERcase"]) 为您提供了一个映射对象你可以遍历得到["this","has","upper","and","lowercase"],因为str.lower被调用了)。
当我们将两者结合起来时,我们将 collections.Counter 提供给 map 对象,该对象将 line.split() 中的每个单词都小写,然后将其添加到用作累加器的初始空的 Counter 中。 Capisce?
破旧不堪
不清楚你的代码有什么问题,所以我只是为你抛出一些知识,希望能坚持下去。
r = input("insert the name of the file")
# this will be a string from the user, containing the file name, e.g.
# r == "text.txt"
# this is normal, because you pass `open` a filename, not a file object
File = open(r, "r")
# this makes File a file object that's pointed at the file name given from
# the user, opened for reading.
data = File.read()
# this sets data equal to the string containing the entire text in File
# This is usually NOT what you want to do, but without further explanation,
# I'll leave it be
data.split()
# this isn't an in-place operation, so you built a list out of the string
# data, split on newlines, then threw it away since you didn't assign it to
# anything.
print(data)
# prints your original data variable, because remember data.split() is not
# in-place, you'd have to do data = data.split(), but that's the wrong way
# to do that anyway....
这就是我认为你想要做的......
filename = input("insert the name of the file: ")
with open(filename, "r") as infile:
data = infile.readlines()
这使用上下文管理器 (with) 而不是 File = open(filename),因为这是一种更好的做法。它基本上使您不必在完成后键入File.close(),并且还说明了在您处理文件时可能会出错的事实,所以如果您的代码出于某种原因抛出异常并且没有't GET to your File.close(),它仍然会在离开 with 块后关闭文件对象。
它还使用.readlines() 而不是.read().split(),这实际上是一回事。这可能仍然不是您想要做的(在大多数情况下,您只想遍历文件而不是将其所有数据转储到内存中)但没有更多上下文,我无法进一步帮助您。
它还遵循 PEP8 的命名约定,其中 Capitalizednames 是类。 File 不是一个类,它是一个文件对象,所以我将它命名为 infile。我通常使用in_ 和out 作为文件名,但YMMV。
如果您评论您要对文件执行的操作,我可以为您编写一些特定的代码。