【问题标题】:how to get inner html properties of a div tag in beautifulsoup如何在beautifulsoup中获取div标签的内部html属性
【发布时间】:2018-01-13 05:41:39
【问题描述】:

网站内建有内部 HTML

美汤不提取嵌入的 HTML 代码。

我需要用 class= qwjRop 提取 div 元素

例如无法从 div 标签中提取“在这个价格上不错”

import requests
from bs4 import BeautifulSoup

url="https://www.flipkart.com/hp-pentium-quad-core-4-gb-1-tb-hdd-dos-15-be010tu-notebook/product-reviews/itmeprzhy4hs4akv?page1&pid=COMEPRZBAPXN2SNF"


def clawler(in_url):
    source_code = requests.get(in_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text, "html.parser")    

    for name in soup.findAll('div',{'class':'qwjRop'}):
       print(name.prettify())

【问题讨论】:

  • 您能给我们一个您在解析时遇到问题的 HTML 示例吗? “嵌入式 HTML 代码”到底是什么意思?你是说 iframe 吗?
  • 编辑了完整的代码,请查看...

标签: python-3.x beautifulsoup web-crawler


【解决方案1】:

当然,我们可以像朋友之前所说的那样使用 Selenium。 这里我要介绍另外一个工具,你可以像Scrapy一样使用它,它叫scrapy_splash,是Scrapy团队创建的Scrapy插件。 使用pip install scrapy_splash并享受它,文档很详细 你可以这样写,scrapy_splash 会为你呈现网站

import scrapy
import scrapy_splash as scrapys
class StaticsSpider(scrapy.Spider):
    name = 'statics'
    start_urls = [
    'https://stackoverflow.com/',
    ]
    def start_requests(self):
        for item in self.start_urls:
            yield scrapys.SplashRequest(
                item, callback=self.parse, args={'wait': 0.5})

    def parse(self, response):
        ......

response会被渲染成website,如果你知道scrapy中如何处理response的话也可以这样使用

【讨论】:

    【解决方案2】:

    页面是用 JavaScript 渲染的,你可以使用 Selenium 来渲染它:

    首先安装 Selenium:

    sudo pip3 install selenium
    

    然后获取驱动程序https://sites.google.com/a/chromium.org/chromedriver/downloads,如果您使用的是 Windows 或 Mac,则可以使用无头版本的 chrome“Chrome Canary”。

    import bs4 as bs
    from selenium import webdriver  
    browser = webdriver.Chrome()
    url="https://www.flipkart.com/hp-pentium-quad-core-4-gb-1-tb-hdd-dos-15-be010tu-notebook/product-reviews/itmeprzhy4hs4akv?page1&pid=COMEPRZBAPXN2SNF"
    browser.get(url)
    html_source = browser.page_source
    browser.quit()
    soup = bs.BeautifulSoup(html_source, "html.parser")
    for name in soup.findAll('div',{'class':'qwjRop'}):
       print(name.prettify())
    

    或者对于其他非硒方法,请参阅我对Scraping Google Finance (BeautifulSoup)的回答

    【讨论】:

    • 非常感谢,从早上开始就挠头解决这个问题。
    猜你喜欢
    • 1970-01-01
    • 2020-11-27
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-03-12
    • 2021-12-23
    相关资源
    最近更新 更多