【问题标题】:Why does this url raise BadStatusLine with httplib2 and urllib2?为什么这个 url 会使用 httplib2 和 urllib2 引发 BadStatusLine?
【发布时间】:2012-02-13 18:10:18
【问题描述】:

使用 httplib2 和 urllib2,我试图从这个 url 获取页面,但所有这些页面都没有成功,最终出现了这个异常。

content = conn.request(uri="http://www.zdnet.co.kr/news/news_print.asp?artice_id=20110727092902")
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/usr/lib/python2.7/dist-packages/httplib2/__init__.py", line 1129, in request
    (response, content) = self._request(conn, authority, uri, request_uri, method, body, headers, redirections, cachekey)
  File "/usr/lib/python2.7/dist-packages/httplib2/__init__.py", line 901, in _request
    (response, content) = self._conn_request(conn, request_uri, method, body, headers)
  File "/usr/lib/python2.7/dist-packages/httplib2/__init__.py", line 871, in _conn_request
    response = conn.getresponse()
  File "/usr/lib/python2.7/httplib.py", line 1027, in getresponse
    response.begin()
  File "/usr/lib/python2.7/httplib.py", line 407, in begin
    version, status, reason = self._read_status()
  File "/usr/lib/python2.7/httplib.py", line 371, in _read_status
    raise BadStatusLine(line)

HTTP头是这样的

http://www.zdnet.co.kr/news/news_print.asp?artice_id=20110727092902

GET /news/news_print.asp?artice_id=20110727092902 HTTP/1.1
Host: www.zdnet.co.kr
User-Agent: Mozilla/5.0 (Windows NT 5.1; rv:10.0.1) Gecko/20100101 Firefox/10.0.1
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8
Accept-Language: ko-kr,ko;q=0.8,en-us;q=0.5,en;q=0.3
Accept-Encoding: gzip, deflate
Connection: keep-alive
Cookie: RMID=7d83495d4f336fe0; __utma=37206251.1552605885.1328771258.1328771258.1329070845.2; __utmz=37206251.1328771258.1.1.utmcsr=(direct)|utmccn=(direct)|utmcmd=(none); ASPSESSIONIDCSQCQTDD=BCLEHPPDEPHEBJDLCFNDMKDN; __utmc=37206251; ASPSESSIONIDSSQCQQCB=MJPLMOJAFPDFCLONCANBIKHN; _EXEN=2
X-FireLogger: 1.2

HTTP/1.1 200 OK
Date: Mon, 13 Feb 2012 18:02:56 GMT
Content-Length: 19158
Content-Type: text/html;charset=UTF-8; Charset=UTF-8
Set-Cookie: ASPSESSIONIDSQSDQRDB=NGAIFHKAGDIOGEMANAOLLKKF; path=/
Cache-Control: private

有什么线索吗?

【问题讨论】:

  • 请发布您的连接声明。

标签: python urllib2 httplib2


【解决方案1】:

这对我来说很好用:

import urllib2

opener = urllib2.build_opener()

headers = {
  'User-Agent': 'Mozilla/5.0 (Windows NT 5.1; rv:10.0.1) Gecko/20100101 Firefox/10.0.1',
}

opener.addheaders = headers.items()
response = opener.open("http://www.zdnet.co.kr/news/news_print.asp?artice_id=20110727092902")

print response.headers
print response.read()

网站会丢弃所有没有User-Agent 字符串的请求。

【讨论】:

    【解决方案2】:

    对于所有在安装 httplib2 0.8 后遇到类似问题的人:

    0.8 版在与 HTTP 保持活动相关的连接处理方面存在回归问题。查看错误报告:https://code.google.com/p/httplib2/issues/detail?id=250

    有针对此问题的修复程序,但目前尚未发布。在此之前只需使用 httplib2 0.7.7。

    【讨论】:

      【解决方案3】:

      在我的代码中,当我使用

          from urllib2 import urlopen  
          content = urlopen(page).read()
      

      出现异常。但是,当我使用

          import urllib  
          content = urllib.urlopen(page).read()
      

      一切正常。 也许它会帮助你。

      【讨论】:

      • 这解决了我的问题。谢谢。但是很高兴知道为什么 urllib2 不起作用。
      【解决方案4】:

      看起来此网页不允许您的用户代理。你可以这样改变它:

      >>> import urllib2
      >>> user_agent = 'Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)'
      >>> headers = { 'User-Agent' : user_agent }
      >>> r = urllib2.Request('http://www.zdnet.co.kr/news/news_print.asp?artice_id=20110727092902', headers=headers)
      >>> fd = urllib2.urlopen(r)
      >>> print fd[20:]
      '<!DOCTYPE html PUBLI'
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2013-10-25
        相关资源
        最近更新 更多