【发布时间】:2020-10-22 15:56:09
【问题描述】:
我正在尝试抓取所有链接到这个网站的形成:https://www.formatic-centre.fr/formation/
所以首先,我启动这个脚本只是为了测试,看看我是否可以抓取第一页的链接:
import scrapy
class LinkSpider(scrapy.Spider):
name = "link"
#allow_domains = ['https://www.formatic-centre.fr/']
start_urls = ['https://www.formatic-centre.fr/formation/']
#rules = (Rule(LinkExtractor(allow=r'formation'), callback="parse", follow= True),)
def parse(self, response):
card = response.xpath('//a[@class="title"]')
for a in card:
yield {'links': a.xpath('@href').get()}
成功了,我得到了这个:
[
{"links": "https://www.formatic-centre.fr/formation/les-regles-juridiques-du-teletravail/"},
{"links": "https://www.formatic-centre.fr/formation/mieux-gerer-son-stress-en-periode-du-covid-19/"},
{"links": "https://www.formatic-centre.fr/formation/dynamiser-vos-equipes-special-post-confinement/"},
{"links": "https://www.formatic-centre.fr/formation/conduire-ses-entretiens-specifique-post-confinement/"},
{"links": "https://www.formatic-centre.fr/formation/cours-excel/"},
{"links": "https://www.formatic-centre.fr/formation/autocad-3d-2/"},
{"links": "https://www.formatic-centre.fr/formation/concevoir-et-developper-une-strategie-marketing/"},
{"links": "https://www.formatic-centre.fr/formation/preparer-soutenance/"},
{"links": "https://www.formatic-centre.fr/formation/mettre-en-place-une-campagne-adwords/"},
{"links": "https://www.formatic-centre.fr/formation/utiliser-google-analytics/"}
]
但是当我想爬取所有页面时,事情变得很脏......我迷路了,我的脚本不再工作,我猜我的循环不太正确,因为我有 doublon 等等。
这是我的最终脚本:
import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from lxml import html
class LinkSpider(scrapy.Spider):
name = "link"
#allow_domains = ['https://www.formatic-centre.fr/']
start_urls = ['https://www.formatic-centre.fr/formation/']
rules = (Rule(LinkExtractor(allow=r'formation'), callback="parse", follow= True),)
def parse(self, response):
card = response.xpath('//a[@class="title"]')
for a in card:
yield {'links': a.xpath('@href').get()}
next_page = response.xpath('.//a[@class="pagination__next btn-squae"]/@href').extract_first()
if next_page:
yield scrapy.Request(
response.urljoin(next_page),
callback=self.parse
)
也许这是我的路?我检查并仔细检查以查看“下一个按钮”在哪里放置好的 href,然后我放了:.//a[@class="pagination__next btn-squae"]/@href
但奇怪的是,html 源代码中的链接没有链接到第二页,所以我很困惑。
这里 -> link
有什么想法吗?
编辑:显然我需要一些FormRequest,我需要使用这种代码吗? ajax
【问题讨论】:
-
你想收集所有内部链接还是特定的东西?
-
一些具体的东西。不是所有的链接,链接者指的是形成。就像上面的例子一样。 (第一张图片)
-
对不起?每个页面的链接?
-
不不,每个格式的链接都像这样:formatic-centre.fr/formation/… 用于所有页面的所有 kinf
标签: python python-3.x web-scraping scrapy