【问题标题】:BeautifulSoup findAll method is not finding all img tags?BeautifulSoup findAll 方法没有找到所有 img 标签?
【发布时间】:2015-03-24 12:52:52
【问题描述】:

我正在写一个网络爬虫脚本

picture = soup.find("div", {"id" : "thumbs-top"})
for link in picture.findAll("img"):
    counter += 1

这个计数器结果是 327,这显然不是真的。我正在刮这个 imgur 画廊:http://imgur.com/a/akHsJ?gallery

【问题讨论】:

    标签: python beautifulsoup


    【解决方案1】:

    我得到 575。

    >>> import requests
    >>> from bs4 import BeautifulSoup
    >>> r = requests.get('http://imgur.com/a/akHsJ?gallery')
    >>> soup = BeautifulSoup(r.text)
    >>> picture = soup.find("div", {"id" : "thumbs-top"})
    >>> links = picture.findAll("img")
    >>> len(links)
    575
    

    编辑:在 cmets 中,我们确定通过使用 BeautifulSoup 尝试不同的 html 解析器来解决问题。有关详细信息,请参阅this 问题。

    >>> soup = BeautifulSoup(r.text, 'html.parser')
    

    【讨论】:

    • 我刚刚在 IDLE 中准确输入了你所拥有的内容,它显示为 327。这是否意味着我的请求或 BeautifulSoup 可能已过时?编辑:我只是通过管道传输了两者,它说它们都已启动迄今为止
    • 您可以尝试打印出 r.text 并尝试查看发生了什么吗?关于浏览器中的页面,我注意到的一件事是您必须向下滚动才能加载更多图像。在我运行我的代码之前,我认为这可能是一个问题,但至少对我来说不是。我怀疑你的库有什么问题。
    • 如果 html 源看起来没问题,然后尝试使用一些不同的解析器和 BeautilfSoup,例如,尝试soup = BeautifulSoup(r.text, 'html.parser')
    • hm... 我打印出 r.text 并且 html 源代码看起来不错。我尝试了 html.parser 并且显然可以工作,因为它打印出 575。你知道为什么它对我不起作用但它对你有用吗?编辑:是的 lxml 解析器(默认的)给出 327 而 html.parser 给出第575章
    • 很可能 html 已损坏(大多数网站并未完全遵循 html 规范)。奇怪的是,我使用的是默认解析器,然而。但是看到这个答案:stackoverflow.com/questions/16322862/…
    猜你喜欢
    • 2017-01-20
    • 2021-02-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-03-06
    • 2016-05-27
    相关资源
    最近更新 更多