【问题标题】:Porting from Python 2 to Python 3: 'utf-8 codec can't decode byte'从 Python 2 移植到 Python 3:“utf-8 编解码器无法解码字节”
【发布时间】:2015-12-15 13:40:32
【问题描述】:

嘿,我尝试将那个小 sn-p 从 2 移植到 Python 3。

Python 2:

def _download_database(self, url):
  try:
    with closing(urllib.urlopen(url)) as u:
      return StringIO(u.read())
  except IOError:
    self.__show_exception(sys.exc_info())
  return None

Python 3:

def _download_database(self, url):
  try:
    with closing(urllib.request.urlopen(url)) as u:
      response = u.read().decode('utf-8')
      return StringIO(response)
  except IOError:
    self.__show_exception(sys.exc_info())
  return None

但我还是得到了

utf-8 codec can't decode byte 0x8f in position 12: invalid start byte

我需要使用 StringIO,因为它是一个 zipfile,我想用那个函数解析它:

   def _parse_zip(self, raw_zip):
  try:
     zip = zipfile.ZipFile(raw_zip)

     filelist = map(lambda x: x.filename, zip.filelist)
     db_file  = 'IpToCountry.csv' if 'IpToCountry.csv' in filelist else filelist[0]

     with closing(StringIO(zip.read(db_file))) as raw_database:
        return_val = self.___parse_database(raw_database)

     if return_val:
        self._load_data()

  except:
     self.__show_exception(sys.exc_info())
     return_val = False

  return return_val

raw_zip 是下载数据库函数的返回

【问题讨论】:

  • 您收到的数据的编码显然是不是 UTF-8。它是什么编码?如果 Web 服务器是正确的,那么 HTTP 响应的 Content-Type 标头应该告诉您,以及文档中可能的 HTML <meta> 标记(如果是 HTML)。
  • 许多网络服务器的默认编码是iso-8859-1。
  • Here 是 StackOverflow 上的现有问题,其答案解释了将字节解码为字符。
  • 那个 url 正在下载一个 zip 文件 -- 你为什么要把一个二进制文件转换成一个字符串?
  • @Fragkiller 正如其他人所指出的,您正在检索二进制文件,而不是文本。 Ashley Wilson 的回答显示了如何获取字节。它在 Python 2 中“工作”的原因是因为 Python 2 在字节与字符方面非常草率,并且没有很好地处理 Unicode。在 Python 3 中,您需要了解其中的区别。

标签: python-3.x urllib stringio


【解决方案1】:

utf-8 无法解码任意二进制数据。

utf-8 是一种字符编码,可用于将文本(例如,在 Python 3 中表示为str 类型——Unicode 代码点序列)编码为字节串(bytes 类型——字节序列([0, 255] 区间内的小整数))并将其解码回来。

utf-8 不是唯一的字符编码。有些字符编码与 utf-8 不兼容。即使.decode('utf-8') 没有引发异常;这并不意味着结果是正确的——如果你使用错误的字符编码来解码文本,你可能会得到mojibake。见A good way to get the charset/encoding of an HTTP response in Python。

您的输入是一个 zip 文件——二进制数据不是文本,因此您不应尝试将其解码为文本。

Python 3 可帮助您查找与混合二进制数据和文本相关的错误。 要将代码从 Python 2 移植到 Python 3,您应该了解文本 (Unicode) 与二进制数据(字节)的区别。

str 在 Python 2 上是一个字节串,可用于二进制数据和(编码)文本。除非from __future__ import unicode_literals 存在; '' 文字在 Python 2 中创建一个字节串。u'' 创建 unicode 实例。在 Python 3 str 上,类型是 Unicode。 bytes 指的是 Python 3 和 Python 2.7 上的字节序列(bytes 是 Python 2 上 str 的别名)。 b'' 在 Python 2/3 上创建 bytes 实例。

urllib.request.urlopen(url) 返回一个类似文件的对象(二进制文件),您可以按原样传递它在某些情况下 例如,to decode remote gzipped content on-the-fly:

#!/usr/bin/env python3
import xml.etree.ElementTree as etree
from gzip import GzipFile
from urllib.request import urlopen, Request

with urlopen(Request("http://smarkets.s3.amazonaws.com/oddsfeed.xml",
                     headers={"Accept-Encoding": "gzip"})) as response, \
     GzipFile(fileobj=response) as xml_file:
    for elem in getelements(xml_file, 'interesting_tag'):
        process(elem)

ZipFile() 需要一个seek()-able 文件,因此您不能直接传递urlopen()。您必须先下载内容。您可以使用io.BytesIO() 来包装它:

#!/usr/bin/env python3
import io
import zipfile
from urllib.request import urlopen

url = "http://www.pythonchallenge.com/pc/def/channel.zip"
with urlopen(url) as r, zipfile.ZipFile(io.BytesIO(r.read())) as archive:
    print({member.filename: archive.read(member) for member in archive.infolist()})

StringIO() 是文本文件。它在 Python 3 中存储 Unicode。

【讨论】:

    【解决方案2】:

    如果您感兴趣的只是从您的函数返回一个流处理程序(而不是要求解码内容),您可以使用BytesIO 而不是StringIO:

    from contextlib import closing
    from io import BytesIO
    from urllib.request import urlopen
    
    url = 'http://www.google.com'
    
    
    with closing(urlopen(url)) as u:
        response = u.read()
        print(BytesIO(response))
    

    【讨论】:

      【解决方案3】:

      您发布的链接http://software77.net/geo-ip?DL=2 正在尝试下载zip 文件,该文件是二进制文件。

      • 不应将二进制 blob 转换为 str(只需使用 BytesIO)
      • 如果您有充分的理由这样做,请使用latin-1 作为解码器。

      【讨论】:

        猜你喜欢
        • 2020-07-17
        • 1970-01-01
        • 2018-05-05
        • 1970-01-01
        • 1970-01-01
        • 2019-11-10
        • 1970-01-01
        • 2023-01-30
        • 2015-08-24
        相关资源
        最近更新 更多