【问题标题】:Not able to store data scraped with scrapy in json or csv format无法以 json 或 csv 格式存储用 scrapy 抓取的数据
【发布时间】:2017-03-06 14:34:53
【问题描述】:

在这里,我想存储网站页面上给出的列表中的数据。如果我正在运行命令

response.css('title::text').extract_first()        and
response.css("article div#section-2 li::text").extract()

单独在scrapy shell中,它在shell中显示了预期的输出。 下面是我的代码,它没有以 json 或 csv 格式存储数据:

import scrapy

class QuotesSpider(scrapy.Spider):
    name = "medical"

    start_urls = ['https://medlineplus.gov/ency/article/000178.html/']


    def parse(self, response):
        yield
        {
            'topic': response.css('title::text').extract_first(),
            'symptoms': response.css("article div#section-2 li::text").extract()
        }

我已尝试使用

运行此代码
scrapy crawl medical -o medical.json

【问题讨论】:

    标签: json csv web-scraping scrapy


    【解决方案1】:

    您需要修正您的网址,它是https://medlineplus.gov/ency/article/000178.htm 而不是https://medlineplus.gov/ency/article/000178.html/。

    此外,更重要的是,您需要定义一个 Item 类并从蜘蛛的 parse() 回调中产生/返回它:

    import scrapy
    
    
    class MyItem(scrapy.Item):
        topic = scrapy.Field()
        symptoms = scrapy.Field()
    
    
    class QuotesSpider(scrapy.Spider):
        name = "medical"
    
        allowed_domains = ['medlineplus.gov']
        start_urls = ['https://medlineplus.gov/ency/article/000178.htm']
    
        def parse(self, response):
            item = MyItem()
    
            item["topic"] = response.css('title::text').extract_first()
            item["symptoms"] = response.css("article div#section-2 li::text").extract()
    
            yield item
    

    【讨论】:

      猜你喜欢
      • 2016-12-11
      • 2020-08-18
      • 2022-08-04
      • 1970-01-01
      • 2012-01-23
      • 2021-10-07
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多