【问题标题】:python utf-8 problempython utf-8 问题
【发布时间】:2011-07-19 15:35:44
【问题描述】:

这是我的脚本

# -*- coding: utf-8 -*-
from BeautifulSoup import BeautifulSoup
import urllib2

res = urllib2.urlopen('http://tazeh.net')
html = res.read()

soup = BeautifulSoup(''.join(html))

title = soup.findAll('title')
print title

当我在终端中运行此脚本时,我会收到这样的错误文本

$ python test.py

[<title>ٞاŰ&OElig;گاŮ&Dagger; ؎بعŰ&OElig; ŘŞŘ­Ů&bdquo;Ű&OElig;Ů&bdquo;Ű&OElig; تازŮ&Dagger;</title>]

此标题采用 utf-8 编码和波斯语

我是python的新手,怎么了?

【问题讨论】:

  • 你试过title.decode()吗?
  • 将脚本底部更改为code title = soup.findAll('title') title = title[0].string.decode('utf-8') print title code got error return codecs.utf_8_decode(input, errors, True) UnicodeEncodeError: 'ascii' codec can't encode characters in position 0-4: ordinal not in range(128)
  • @Efazati 这不是你的事:D

标签: python utf-8 persian


【解决方案1】:

如果我添加(就像建议在不太有用的地方做的 cmets 之一):

html = html[:10000].decode("utf-8")

(切片是因为解码在页面更远的偏移处失败)

之前:

soup = BeautifulSoup(html)

打印出来:

[<title>پایگاه خبری تحلیلی تازه</title>]

【讨论】:

  • 切片 [:10000] 是因为在页面更远的偏移处解码失败。
【解决方案2】:

''.join(html) 是不必要的。变量html 已经是一个字符串。

但是,页面似乎没有正确地以 UTF-8 编码。

【讨论】:

    猜你喜欢
    • 2017-04-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多