【问题标题】:How to scrape the items loaded via a "view more" button using Scrapy如何使用 Scrapy 抓取通过“查看更多”按钮加载的项目
【发布时间】:2023-04-04 13:59:01
【问题描述】:

这里是网站中查看更多按钮的检查。我可以抓取网站中显示的数据,但我希望它可以抓取隐藏在“查看更多”按钮后面的项目。我怎么做?

 <div id="view-more" class="p20px pt10px">
                        <div id="view-more-loader" class="tac"></div>

                        <a href="javascript:void(0);" onclick="add_more_product_classified();$('#load_more_a_id').hide();" class="xxxxlarge ffrc lightbginfo gbiwb bdr darkbdrinfo p10px20px db w180px m0a tac" id="load_more_a_id" style="display: block;"><b class="icon-refresh xsmall mr5px"></b>View More Products..</a>
                        </div>

我的scrapy代码:

import scrapy




class DummymartSpider(scrapy.Spider):
    name = 'dummymart'
    allowed_domains = ['dummymart.net']
    start_urls =['https://www.dummymart.com/catalog/car-dvd-player_cid100001018.html']



    def parse(self, response):
            Product = response.xpath('//div[@class="attr"]/h2/a/@title').extract()
            Company =  response.xpath('//div[@class="supplier"]/p/a/@title').extract()
            Country =  response.xpath('//*[@class="location a-color-secondary"]/span/text()').extract()
            Category = response.xpath('//*[@class="attr category hide--mobile"]/span/a/text()').extract()

            for item in zip(Product,Company,Country,Category):
                scraped_info = {
                    'Product':item[0],
                    'Company': item[1],
                    'Country':item[2],
                    'Category':item[3]

                }
                yield scraped_info

【问题讨论】:

    标签: python xpath web-scraping scrapy


    【解决方案1】:

    此类问题的通常解决方案是:

    1. 在浏览器中启动开发者工具;
    2. 转到网络面板,以便查看浏览器发出的请求;
    3. 点击页面中的“查看更多”按钮,查看您的浏览器执行了哪个请求来获取数据;
    4. 对您的蜘蛛发出相同的请求。

    This blog post 可以帮到你。

    【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2018-04-16
    • 1970-01-01
    • 2021-08-25
    • 1970-01-01
    • 2021-09-26
    • 1970-01-01
    • 2019-01-20
    • 2018-07-06
    相关资源
    最近更新 更多