【发布时间】:2015-12-14 16:08:24
【问题描述】:
我是scrapy的新手。 我想在this 页面上抓取产品。我的代码只抓取了第一页,也抓取了大约 15 个产品,然后它就停止了。并且也想抓取下一页。有什么帮助吗?
这是我的课
class AllyouneedSpider(CrawlSpider):
name = "allyouneed"
allowed_domains = ["de.allyouneed.com"]
start_urls = [ 'http://de.allyouneed.com/de/sportschuhe-/8799665488014/',]
rules = (
Rule(LxmlLinkExtractor(allow=(), restrict_xpaths='//*[@class="itm fst jf-lDiv"]//a[@href]'), callback='parse_obj', process_links="parse_filter") ,
Rule(LxmlLinkExtractor(restrict_xpaths='//*[@id="M62_searchhit"]//a[@href]')),
)
def parse_filter(self, links):
for link in links:
if self.allowed_domains[0] not in link.url:
pass # print link.url
# print links
return links
def parse_obj(self, response):
item = AllyouneedItem()
sel = scrapy.Selector(response)
item['url'] = []
url = response.selector.xpath('//*[@id="M62_searchhit"]//a[@href]').extract()
ti = response.selector.xpath('//span[@itemprop="name"]/text()').extract()
dec = response.selector.xpath('//div[@class="m-desc m-desc-t"]//text()').extract()
cat = response.selector.xpath('//span[@itemprop="title"]/text()').extract()
if ti:
item['title'] = ti
item['url'] = response.url
item['category'] = cat
item['decription'] = dec
print item
yield item
【问题讨论】:
-
你能分享日志吗?
-
也请更正缩进。
标签: python scrapy scrapy-spider