【问题标题】:How do I web scrape with beautiful soup an element without class or id我如何用漂亮的汤抓取一个没有类或 id 的元素
【发布时间】:2019-08-29 16:11:06
【问题描述】:

我正在尝试使用漂亮的汤在此网站http://www.singaporepools.com.sg/en/product/Pages/toto_results.aspx 上为“下一个大奖”刮取金额,但我很难获得 1,000,000(目前)

我目前正在阅读这本书:用 python 自动化无聊的东西。我在这里花了将近 2 个小时阅读在线教程和过去的问题,但我仍然无法弄清楚如何为大多数教程显示的没有 class 或 id 的元素执行此操作

import requests, bs4
res=requests.get('http://www.singaporepools.com.sg/en/product/Pages/toto_results.aspx')
res.raise_for_status()
noStarchSoup = bs4.BeautifulSoup(res.text, "html.parser")
elems = noStarchSoup.find('span', {'style':'color'}, {'style':'font-weight'})

【问题讨论】:

    标签: python web-scraping beautifulsoup


    【解决方案1】:

    如果您转到网络选项卡,您会在下面找到网址。

    http://www.singaporepools.com.sg/DataFileArchive/Lottery/Output/toto_next_draw_estimate_en.html?v=2019y8m29d17h15m

    使用以下代码检索值。

    import requests, bs4
    import re
    res=requests.get('http://www.singaporepools.com.sg/DataFileArchive/Lottery/Output/toto_next_draw_estimate_en.html?v=2019y8m29d17h15m')
    res.raise_for_status()
    noStarchSoup = bs4.BeautifulSoup(res.text, "html.parser")
    elems = noStarchSoup.find('div',text=re.compile('Next Jackpot')).find_next('span')
    print(elems.text)
    

    输出:

    $1,000,000 est
    

    【讨论】:

    • 我想这会起作用,但为了获得奖励,我们应该指出查询参数 v 是硬编码的。 OP 应该按日期构建这个字符串
    • @RafaelT : JavaScripts 渲染到页面。我相信 selenium 和 bs4 应该可以正常工作。
    • 我的意思是参数值是v=2019y8m29d17h15m,这对我来说就像一个时间戳。 OP 应该构建它以获取当前值,而不是每天都一样
    • 感谢您的回答,在进行网络抓取以隔离/减少您需要查看的 html 时检查网络选项卡是一个好习惯吗?这里是初学者。
    • 是的,如果它是动态页面,那么您需要检查网络选项卡并查找 API
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-07-30
    • 1970-01-01
    • 2015-01-26
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多