【问题标题】:Python-read a file from a defined variable [closed]Python-从定义的变量中读取文件[关闭]
【发布时间】:2014-03-28 17:37:22
【问题描述】:

我想让用户 input 一个要在 Python 中读取的文件名(例如:text.txt),但它读取为字符串而不是文件类型。

r=(input("insert the name of the file"))
  File= open(r,'r')
  data=File.read()
  data.split()
  print(data)

【问题讨论】:

  • 等等,什么?我完全不明白你在问什么。当然它读为一个字符串,为什么它会做其他事情呢?
  • 您的代码现在完全符合您的要求。你遇到了什么问题?
  • 而不是文件类型:您期望根据文件类型做什么?
  • 也许你想读
  • 缩进有点偏离,并且有一对额外的 () 是不需要的。除此之外,代码是正确的。

标签: python file input


【解决方案1】:

新热点

编辑:根据我的回答中的 cmets,OP 正在寻求为文件中的所有单词(空格分隔)构建一个包含 {word:wordcount} 的 dict。

有一个非常棒的方法可以做到这一点,但它并没有真正教给你任何东西,所以我会先向你展示缓慢的方法,然后再包括最佳解决方案。

wordcountdict = dict()

r = input("filename: ")
with open(r, 'r') as infile:
    for line in infile:
        for word in infile.split(): # split on whitespace
            try:
                wordcountdict[word.lower()] += 1
                # try adding one to the word in the counter
            except KeyError:
                wordcountdict[word.lower()] = 1
                # If the word isn't in the dict already, set it to 1

现在您可能想要过滤掉一些常用词("at"、"I"、"then" 等),在这种情况下,您可以建立它们的黑名单(例如 blacklist = ['at', 'i', 'then'])并执行 if word.lower() in blacklist: continue在for word in infile.split() 内部和try/except 块之前。这将测试该单词是否在黑名单中,如果是则跳过该执行的其余部分。

现在我向你保证了一个很好的方法来做到这一点,那就是collections.Counter。它是专门为计算列表中的元素而创建的字典。有更快的方法来计算单词,但在 Python (imo) 中没有更干净的方法。你在this question查看时间安排

from collections import Counter

wordcountdict = Counter()
r = input("filename: ")
with open(r, 'r') as infile:
    for line in infile:
        wordcountdict += Counter( map(str.lower,line.split()) )

如果您从未使用过来自 collections 或 map 函数的导入,那么这将是非常神秘的,这就是我没有把它放在首位的原因! :)。

基本上:collections.Counter 将一个可迭代对象作为参数,并计算可迭代对象中的所有元素(因此 `Counter([1,1,2,3,4,4,4]) == {1:2, 2:1、3:1、4:3})。您可以添加它们,它会在它们唯一的地方创建新的键,并在它们不唯一的地方添加值。

map(callable, iterable) 运行 callable 并带有可迭代的每个元素的参数,并返回一个本身可迭代的 map 对象(在 Python2 中是 list)(因此 map(str.lower, ["ThIS", "Has", "UppEr", "aNd", "LOWERcase"]) 为您提供了一个映射对象你可以遍历得到["this","has","upper","and","lowercase"],因为str.lower被调用了)。

当我们将两者结合起来时,我们将 collections.Counter 提供给 map 对象,该对象将 line.split() 中的每个单词都小写,然后将其添加到用作累加器的初始空的 Counter 中。 Capisce?

破旧不堪

不清楚你的代码有什么问题,所以我只是为你抛出一些知识,希望能坚持下去。

r = input("insert the name of the file")
# this will be a string from the user, containing the file name, e.g.
# r == "text.txt"
# this is normal, because you pass `open` a filename, not a file object

File = open(r, "r")
# this makes File a file object that's pointed at the file name given from
# the user, opened for reading.

data = File.read()
# this sets data equal to the string containing the entire text in File
# This is usually NOT what you want to do, but without further explanation,
# I'll leave it be

data.split()
# this isn't an in-place operation, so you built a list out of the string
# data, split on newlines, then threw it away since you didn't assign it to
# anything.

print(data)
# prints your original data variable, because remember data.split() is not
# in-place, you'd have to do data = data.split(), but that's the wrong way
# to do that anyway....

这就是我认为你想要做的......

filename = input("insert the name of the file: ")
with open(filename, "r") as infile:
    data = infile.readlines()

这使用上下文管理器 (with) 而不是 File = open(filename),因为这是一种更好的做法。它基本上使您不必在完成后键入File.close(),并且还说明了在您处理文件时可能会出错的事实,所以如果您的代码出于某种原因抛出异常并且没有't GET to your File.close(),它仍然会在离开 with 块后关闭文件对象。

它还使用.readlines() 而不是.read().split(),这实际上是一回事。这可能仍然不是您想要做的(在大多数情况下,您只想遍历文件而不是将其所有数据转储到内存中)但没有更多上下文,我无法进一步帮助您。

它还遵循 PEP8 的命名约定,其中 Capitalizednames 是类。 File 不是一个类,它是一个文件对象,所以我将它命名为 infile。我通常使用in_ 和out 作为文件名,但YMMV。

如果您评论您要对文件执行的操作,我可以为您编写一些特定的代码。

【讨论】:

  • 最好使用read().splitlines(),因为readlines() 留下了Python 在扫描文件时得到的烦人的换行符'\n'。
  • @AlexThornton list(map(strip,map(splitlines,''.join([char for char in infile.read()])))) 以获得更多混淆 :)
  • 很好的解释!但我想做的是接收一个文件并计算该文件中的单词数,然后将它们放入带有{word,numberoftimes}的字典中。我是 python 新手,所以我还在学习
  • @user3448350 我会编辑,但你应该把它放在问题中,这样我们才能重新打开。
【解决方案2】:

我不介意在黑暗中拍摄。但是,如果您发布要阅读的文件之一,这将有所帮助。并且您需要考虑该人现在知道他们应该输入哪个文件,以及当他们不输入您认为他们会输入的内容时会发生什么。

r = raw_input('type the name of the file: ')
with open(r,'r') as myfile:
    for data in myfile:
        print(data.split())

【讨论】:

  • 这里有一个问题,OP 代码中的data 并不意味着您的代码中的data。您的 data 只是文件的一行。无论如何都不错,只是可能会让 OP 感到困惑
  • 好点。它对如何使用/格式化数据更有指导意义。
猜你喜欢
  • 2014-06-30
  • 1970-01-01
  • 1970-01-01
  • 2021-05-20
  • 2016-01-11
  • 1970-01-01
  • 2018-08-04
  • 1970-01-01
  • 2013-03-13
相关资源
最近更新 更多