【问题标题】:Encoding (UTF-8) issue编码 (UTF-8) 问题
【发布时间】:2019-04-21 12:23:34
【问题描述】:

我想从列表中写入文本。但是编码不工作和像位一样写入。

with open('freq.txt', 'w') as f:
    for item in freq:
        f.write("%s\n" % item.encode("utf-8"))

输出:

b'okul'
b'y\xc4\xb1l\xc4\xb1'

预期:

okul
yılı

【问题讨论】:

    标签: python python-3.x utf-8 character-encoding


    【解决方案1】:

    如果您使用的是 Python3,您可以在对 open 的调用中声明您想要的编码:

    with open('freq.txt', 'w', encoding='utf-8') as f:
        for item in freq:
            f.write("%s\n" % item)
    

    如果您不提供编码,则默认为locale.getpreferredencoding() 返回的编码。

    您的代码的问题是'%s\n' % item.encode('utf-8') 将item 编码为字节,但是字符串格式化操作隐式调用字节上的str,这导致字节的repr 被用于构造字符串。

    >>> s = 'yılı'
    >>> bs = s.encode('utf-8')
    >>> bs
    b'y\xc4\xb1l\xc4\xb1'
    >>> # See how the "b" is *inside* the string.
    >>> '%s' % bs
    "b'y\\xc4\\xb1l\\xc4\\xb1'"
    

    将格式字符串设为bytes 文字可避免此问题

    >>> b'%s' % bs
    b'y\xc4\xb1l\xc4\xb1'
    

    但随后写入文件会失败,因为您无法将字节写入以文本模式打开的文件。如果你真的想手动编码,你必须这样做:

    # Open the file in binary mode.
    with open('freq.txt', 'wb') as f:
        for item in freq:
            # Encode the entire string before writing to the file.
            f.write(("%s\n" % item).encode('utf-8'))
    

    【讨论】:

      【解决方案2】:
      import codecs
      
      with codecs.open("lol", "w", "utf-8") as file:
          file.write('Okul')
          file.write('yılı')
      

      【讨论】:

        猜你喜欢
        • 2010-12-01
        • 1970-01-01
        • 2013-11-12
        • 2017-12-22
        • 2011-11-08
        • 2013-09-01
        相关资源
        最近更新 更多