【问题标题】:Web Scraping - I cannot use a for loop to list elementWeb Scraping - 我不能使用 for 循环来列出元素
【发布时间】:2019-06-26 21:03:11
【问题描述】:

我目前正在构建一个网络爬虫,但遇到了一个问题。 当我尝试构建我的 for 循环以便按公司重新组合所有信息时,提取会继续显示相同类型的所有元素。

当我意识到它不起作用时,我返回并尝试仅显示第一个元素的索引列表,但即使我键入 [0],所有元素都会显示给我,好像没有进行特定选择

import scrapy
from centech.items import CentechItem

class CentechSpiderSpider(scrapy.Spider):
    name = 'centech_spider'
    start_urls = ['https://centech.co/nos-entreprises/']

    def parse(self, response):
       items = CentechItem()
       all_companies = response.xpath("//div[@class = 'fl-post-carousel- 
    post']")[1]    #   "//div[@class = 'fl-post-carousel-post']")[1]
    Nom = all_companies.xpath("//h2[contains(@class, 'fl-post-carousel- 
    title')]/text()").extract()
    Description = all_companies.xpath("//div[contains(@class, 
    'description')]/p/text()").extract()
    # Nom = all_companies.response.css("h2.fl-post-carousel- 
    title::text").extract()
    # Description = all_companies.xpath("p::text").extract()

    yield {'Nom' : Nom ,
           'Description' : Description ,
           }

我希望只看到页面的第一个元素,但会显示所有企业。

谢谢。

【问题讨论】:

  • 你也可以在 xpath 中添加 id 来唯一标识它

标签: python for-loop web-scraping scrapy


【解决方案1】:

我不太确定您希望获得的输出。我猜测并修改了您的脚本以获取以下结果。你需要深入一层来获取完整的描述,因为一些描述被破坏了:

import scrapy

class CentechSpiderSpider(scrapy.Spider):
    name = 'centech_spider'
    start_urls = ['https://centech.co/nos-entreprises/']

    def parse(self, response):
        for item in response.css("a.fl-post-carousel-link"):
            nom = item.css(".description > h2.fl-post-carousel-title::text").get()
            description = item.css(".description > p::text").get()
            yield {'nom':nom,'description':description}

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-01-26
    • 2014-10-02
    • 2020-03-06
    • 1970-01-01
    相关资源
    最近更新 更多