【发布时间】:2019-10-29 19:37:17
【问题描述】:
我已经制作了一个爬虫,它可以访问这个网站https://www.cartoon3rbi.net/cats.html然后按照第一条规则打开每个节目的链接,通过 parse_title 方法获取它的标题,然后按照第三条规则打开每一集的链接并获取它的名称。它工作正常,我只需要知道如何为每个节目的剧集名称制作一个单独的 csv 文件,并将 parse_title 方法中的标题用作 csv 文件的名称。有什么建议吗?
# -*- coding: utf-8 -*-
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
class FfySpider(CrawlSpider):
custom_settings = {
'CONCURRENT_REQUESTS': 1
}
name = 'FFy'
allowed_domains = ['cartoon3rbi.net']
start_urls = ['https://www.cartoon3rbi.net/cats.html']
rules = (
Rule(LinkExtractor(restrict_xpaths='//div[@class="pagination"]/a[last()]'), follow=True),
Rule(LinkExtractor(restrict_xpaths='//div[@class="cartoon_cat"]'), callback='title_parse', follow=True),
Rule(LinkExtractor(restrict_xpaths='//div[@class="cartoon_eps_name"]'), callback='parse_item', follow=True),
)
def title_parse(self, response):
title = response.xpath('//div[@class="sidebar_title"][1]/text()').extract()
def parse_item(self, response):
for el in response.xpath('//div[@id="topme"]'):
yield {
'name': el.xpath('//div[@class="block_title"]/text()').extract_first()
}
【问题讨论】:
-
这能回答你的问题吗? Export scrapy items to different files
标签: python csv web-scraping scrapy