【问题标题】:How to get a webpage with unicode chars in python如何在python中获取带有unicode字符的网页
【发布时间】:2017-08-23 17:08:49
【问题描述】:

我正在尝试获取并解析包含非 ASCII 字符的网页(URL 为 http://www.one.co.il)。这就是我所拥有的:

url = "http://www.one.co.il"
req = urllib2.Request(url)
response = urllib2.urlopen(req)
encoding = response.headers.getparam('charset') # windows-1255
html = response.read() # The length of this is valid - about 31000-32000,
                       # but printing the first characters shows garbage -
                       # '\x1f\x8b\x08\x00\x00\x00\x00\x00', instead of
                       # '<!DOCTYPE'
html_decoded = html.decode(encoding)

最后一行给了我一个例外:

File "C:/Users/....\WebGetter.py", line 16, in get_page
  html_decoded = html.decode(encoding)
File "C:\Python27\lib\encodings\cp1255.py", line 15, in decode
  return codecs.charmap_decode(input,errors,decoding_table)
UnicodeDecodeError: 'charmap' codec can't decode byte 0xdb in position 14: character maps to <undefined>

我尝试查看其他相关问题,例如 urllib2 read to Unicode 和 How to handle response encoding from urllib.request.urlopen() ,但没有发现任何有用的信息。

有人可以在这个主题上给我一些启发并指导我吗?谢谢!

【问题讨论】:

    标签: python python-2.7 encoding urllib2 windows-1255


    【解决方案1】:

    0x1f 0x8b 0x08 是 gzip 压缩文件的幻数。您需要先将其解压缩,然后才能使用其中的内容。

    【讨论】:

    • 我是否应该在回复中寻找其他类似的“惊喜”?有没有办法透明地获取包含所有需要的后处理的页面,以便我在 Chrome 的视图源中看到它?
    • 我确定有人已经处理过了。环顾四周。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-05-13
    • 1970-01-01
    • 2015-11-16
    • 2021-08-13
    相关资源
    最近更新 更多