【问题标题】:beautifulsoup parsing error in python -- junk characterspython中的beautifulsoup解析错误——垃圾字符
【发布时间】:2014-05-18 22:02:32
【问题描述】:

代码 - 不知道我做了什么让 BeautifulSoup (BS) 不起作用

import mechanize
import urllib2
from bs4 import BeautifulSoup

#create a browser object to login
browser = mechanize.Browser()

#tell the browser we are human, and not a robot, so the mechanize library doesn't block us
browser.set_handle_robots(False)

browser.addheaders = [('User-Agent','Mozilla/5.0 (Windows U; Windows NT 6.0; en-US; rv:9.0.6')]
#url
url = 'https://www.google.com.au/search?q=python'
#open the url in our virtual browser
browser.open(url)
html = browser.response().read()
print html
soup = BeautifulSoup(html)
print(soup.prettify())

错误

HTMLParseError: junk characters in start tag: u'{t:1}); class="gbzt ', at line 1, column 42892

<!doctype html><html itemscope="" itemtype="http://schema.org/WebPage" lang="en-AU"><head><meta content="text/html; charset=UTF-8" http-equiv="Content-Type"><meta content="/images/google_favicon_128.png" itemprop="image"><title>python - Google Search</title><style>#gb{font:13px/27px Arial,sans-serif;height:30px}#gbz,#gbg{position:absolute;white-space:nowrap;top:0;height:30px;z-index:1000}#gbz{left:0;padding-left:4px}#gbg{right:0;padding-right:5px}#gbs{background:transparent;position:absolute;top:-999px;v

【问题讨论】:

  • 您似乎遇到了错误,因为您在 html 中引入 css 并试图将其解析为 HTML。这里有一个类似的问题可能会帮助stackoverflow.com/questions/10401110/…
  • @yoshiserry 代码对我来说运行良好,你使用的是什么版本的 python?
  • 2.7?我应该安装 lxml 解析器吗?也许?
  • 也许是获取 css 和 html 的 mechanize 命令,而 beautofulsoup(bs) 通常只会获取 html?
  • @yoshiserry,我正在使用 2.7 也没有问题,您的脚本获取并打印一切正常。您只想解析页面吗?

标签: python beautifulsoup mechanize


【解决方案1】:

尝试使用requests:

import requests
from bs4 import BeautifulSoup
#url
url = 'https://www.google.com.au/search?q=python'
r=requests.get(url)
html = r.text
print html
soup = BeautifulSoup(html)
print(soup.prettify())

【讨论】:

  • 哇非常感谢它的工作原理。我很困惑为什么beautifulsoup 不起作用。
  • 它确实使用了请求,但它本身没有使用 beautifulsoup;它显然会吐出我下载的垃圾字符。
  • 也许我没有正确安装它,我从 chris 的二进制包中安装的一些其他模块也无法正常工作。我在看着你pyQt4
  • 请求库是否有能力强制页面等到全部下载完毕(javascript动态生成一些页面内容)我想抓取javascript生成的动态内容。
猜你喜欢
  • 2021-12-04
  • 1970-01-01
  • 2022-01-15
  • 1970-01-01
  • 1970-01-01
  • 2018-02-01
  • 2020-04-08
  • 1970-01-01
  • 2010-12-06
相关资源
最近更新 更多