【发布时间】:2015-12-31 20:43:17
【问题描述】:
您好,我正在尝试抓取一个使用滚动加载数据并加载更多按钮的电子商务网站想要一些帮助让我开始我对网络抓取很陌生。
编辑问题 好的,我正在抓取这个网站 [link]http://www.jabong.com/women/,它有子类别,我正在尝试抓取我在上面代码中尝试过的所有子类别产品,但没有奏效,所以在做了一些研究之后,我创建了一个有效但不满足我的目标的代码.到目前为止我已经尝试过这个
` import scrapy
#from scrapy.exceptions import CloseSpider
from scrapy.spiders import Spider
#from scrapy.contrib.spiders import CrawlSpider, Rule
from scrapy.http import Request
from koovs.items import product
#from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor
from scrapy.selector import Selector
class propubSpider(scrapy.Spider):
name = 'koovs'
allowed_domains = ['jabong.com']
max_pages = 40
def start_requests(self):
for i in range(self.max_pages):
yield scrapy.Request('http://www.jabong.com/women/clothing/tops-tees-shirts/tops/?page=%d' % i, callback=self.parse)
def parse(self, response):
for sel in response.xpath('//*[@id="catalog-product"]/section[2]'):
item = product()
item['price'] = sel.xpath('//*[@class="price"]/span[2]/text()').extract()
item['image'] = sel.xpath('//*[@class="primary-image thumb loaded"]/img').extract()
item['title'] = sel.xpath('//*[@data-original-href]/@href').extract()
`
上面的代码适用于一个类别,如果我指定页数,上面的网站有很多给定类别的产品,我不知道它们驻留在多少页,所以我决定使用爬虫来浏览所有类别和产品页面并获取数据,但我对scrapy非常陌生,任何帮助将不胜感激
【问题讨论】:
-
@Alecxe 你能帮我按照你的代码,但它不工作
-
如果您发布您想要抓取的网站以及您尝试过的代码,那么获得帮助会容易得多。
-
@stummjr 我已经添加了我尝试过的代码,并且还按照建议提到了该网站。
-
你能解决问题还是需要帮助?
-
@Rahul 是的,我仍然需要帮助,我正在耐心等待有人回答
标签: python-2.7 web-scraping scrapy