【问题标题】:Reading Chinese text from a file and printing it to the shell从文件中读取中文文本并打印到shell
【发布时间】:2019-03-24 08:03:53
【问题描述】:

我正在尝试制作一个程序,该程序可以从 .txt 文件中读取行汉字并将它们打印到 Python shell(IDLE?)。

我遇到的问题是尝试对 utf-8 中的字符进行编码和解码,使其实际可以用中文打印。

到目前为止,我有这个:

  file_name = input("Enter the core name of the text you wish to analyze:")+'.txt'

  file = open(file_name, encoding="utf8")

  file = file.read().decode('utf-8').split()

  print(file)

但是,每次我运行代码时,我都会收到此错误提示。

    file = file.read().decode('utf-8').split()
AttributeError: 'str' object has no attribute 'decode'

现在,我不完全确定这意味着什么,因为我是编程语言的新手,所以我想知道是否可以从你们那里得到一些提示。非常感谢!

【问题讨论】:

    标签: python utf-8 foreign-keys decode encode


    【解决方案1】:

    根据您的错误消息,我怀疑 .read() 的输出已经是一个字符串(更准确地说,如果您使用的是 Python 3,则为 unicode 字符点)。

    您是否在没有.decode() 呼叫的情况下尝试过?


    为了更好地处理文件,请使用with 上下文,因为这样可以确保您的文件在退出块后正确关闭。 此外,您可以使用 for line in f 语句遍历文件中的行。

    file_name = input("Enter the core name of the text you wish to analyze:")
    
    with open(file_name + '.txt', encoding='utf8') as f:
        for line in f:
            line = line.strip()   # removes new lines or spaces at the start/end
            print(line)
    

    【讨论】:

      【解决方案2】:

      当您读取在 Python 3 中以这种方式打开的文件时:

      file = open(file_name, encoding="utf8")

      您告诉它该文件以 UTF-8 编码,Python 将自动对其进行解码。 file.read() 已经是一个 Unicode 字符串(Python 3 中的 str 类型),所以你不能再次解码它。只需执行以下操作(不要覆盖 file...这是您的文件句柄):

      data = file.read().split()
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2023-03-21
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2014-03-29
        • 2021-04-02
        相关资源
        最近更新 更多