【问题标题】:Web Scraping contents of ::before ::after CSS Psuedo element using BeautifulSoup使用 BeautifulSoup 抓取 ::before ::after CSS 伪元素的网页内容
【发布时间】:2018-10-22 19:47:45
【问题描述】:

我正在学习网页抓取。我想知道我们如何从下面的元素中获取参与者数量?

<li class="header-hero__stat header-hero__stat--participants">
   ::before
   "255,590 Participants"
   ::after
</li>

我试过的代码

soupy = bs(html,'lxml') 
ul = soupy.find('li',{'class':"header-hero__stats"})

返回None

Target page

【问题讨论】:

    标签: css python-3.x web-scraping beautifulsoup


    【解决方案1】:

    这不是伪元素的内容,而是li节点的文本内容,所以

    li = soup.find('li',{'class':"header-hero__stat--participants"}).text
    

    应该足够提取'255,601 Participants'

    使用.text.split()[0] 仅获取号码

    【讨论】:

    • 谢谢哥们!你能说出为什么下面的代码返回 None 对象吗? soupy=bs(html,'lxml') ul=soupy.find('li',{'class':"header-hero__stats"})
    • 此数据可能由 JavaScript 动态生成,因此您可能需要等待一段时间。你能分享你试过的确切代码吗?
    • @MasooriMenon ,抱歉,出于某种原因,我认为您正在使用 Selenium :) 所以提供的解决方案与 Selenium 有关...该数据可能来自 XHR,因此您需要模拟正确的调用到 API。你能分享目标页面的网址吗?
    • https://www.datacamp.com/courses/intermediate-python-for-data-science 是链接,我正在使用 bs4。再次感谢!
    • @MasooriMenon ,很简单,你使用了错误的选择器。尝试ul=soup.find('ul',{'class':"header-hero__stats"}) 定位ul 节点或li=soup.find('li',{'class':"header-hero__stat--participants"}) 定位目标li 节点
    猜你喜欢
    • 1970-01-01
    • 2021-07-28
    • 2012-12-17
    • 2013-01-09
    • 2020-09-20
    • 2022-12-10
    • 2015-05-06
    • 1970-01-01
    相关资源
    最近更新 更多