【问题标题】:.json export formating in ScrapyScrapy 中的 .json 导出格式
【发布时间】:2018-09-13 02:43:10
【问题描述】:

关于 Scrapy 中 json 导出格式的快速问题。我导出的文件如下所示。

{"pages": {"title": "x", "text": "x", "tags": "x", "url": "x"}}
{"pages": {"title": "x", "text": "x", "tags": "x", "url": "x"}}
{"pages": {"title": "x", "text": "x", "tags": "x", "url": "x"}}

但我希望它采用这种精确的格式。不知何故,我需要在“页面”下获取所有其他信息。

{"pages": [
     {"title": "x", "text": "x", "tags": "x", "url": "x"},
     {"title": "x", "text": "x", "tags": "x", "url": "x"},
     {"title": "x", "text": "x", "tags": "x", "url": "x"}
]}

我在scrapy 或python 方面不是很有经验,但除了导出格式之外,我已经在我的蜘蛛中完成了所有其他工作。这是我刚刚开始工作的 pipelines.py。

from scrapy.exporters import JsonItemExporter
import json

class RautahakuPipeline(object):

    def open_spider(self, spider):
        self.file = open('items.json', 'w')

    def close_spider(self, spider):
        self.file.close()

    def process_item(self, item, spider):
        line = json.dumps(dict(item)) + "\n"
        self.file.write(line)
        return item

这些是我需要提取的 spider.py 中的项目

        items = []
        for title, text, tags, url in zip(product_title, product_text, product_tags, product_url):
            item = TechbbsItem()
            item['pages'] = {}
            item['pages']['title'] = title
            item['pages']['text'] = text
            item['pages']['tags'] = tags
            item['pages']['url'] = url
            items.append(item)
        return items

非常感谢任何帮助,因为这是我项目中的最后一个障碍。

编辑

items = {'pages':[{'title':title,'text':text,'tags':tags,'url':url} for title, text, tags, url in zip(product_title, product_text, product_tags, product_url)]}

这会以这种格式提取 .json

{"pages": [{"title": "x", "text": "x", "tags": "x", "url": "x"}]} {"pages": [{"title": "x", "text": "x", "tags": "x", "url": "x"}]} {"pages": [{"title": "x", "text": "x", "tags": "x", "url": "x"}]}

这越来越好,但我仍然需要文件开头的一个“页面”以及它下面的数组中的所有其他内容。

编辑 2

我认为我的 spider.py 是“pages”被添加到 .json 文件中每一行的原因,我最初应该发布它的整个代码。在这里。

# -*- coding: utf-8 -*-
import scrapy
from urllib.parse import urljoin

class TechbbsItem(scrapy.Item):
    pages = scrapy.Field()
    title = scrapy.Field()
    text= scrapy.Field()
    tags= scrapy.Field()
    url = scrapy.Field()

class TechbbsSpider(scrapy.Spider):
    name = 'techbbs'
    allowed_domains = ['bbs.io-tech.fi']
    start_urls = ['https://bbs.io-tech.fi/forums/prosessorit-emolevyt-ja-muistit.73/?prefix_id=1' #This is a list page full of used pc-part listings
             ]
    def parse(self, response): #This visits product links in the product list page
        links = response.css('a.PreviewTooltip::attr(href)').extract()
        for l in links:
            url = response.urljoin(l)
            yield scrapy.Request(url, callback=self.parse_product)
        next_page_url = response.xpath('//a[contains(.,"Seuraava ")]/@href').extract_first()
        if next_page_url:
           next_page_url =  response.urljoin(next_page_url) 
           yield scrapy.Request(url=next_page_url, callback=self.parse)

    def parse_product(self, response): #This extracts data from inside the links
        product_title = response.xpath('normalize-space(//h1/span/following-sibling::text())').extract()
        product_text = response.xpath('//b[contains(.,"Hinta:")]/following-sibling::text()[1]').re('([0-9]+)')
        tags = "tags" #This is just a placeholder
        product_tags = tags
        product_url = response.xpath('//html/head/link[7]/@href').extract()

        items = []
        for title, text, tags, url in zip(product_title, product_text, product_tags, product_url):
            item = TechbbsItem()
            item['pages'] = {}
            item['pages']['title'] = title
            item['pages']['text'] = text
            item['pages']['tags'] = tags
            item['pages']['url'] = url
            items.append(item)
        return items

所以我的蜘蛛开始从一个充满产品列表的页面爬行。它访问 50 个产品链接中的每一个,并抓取 4 个项目、标题、文本、标签和 url。在抓取一页中的每个链接后,它会转到下一个,依此类推。我怀疑代码中的循环会阻止您的建议对我有用。

我想将 .json 导出为原始问题中提到的确切形式。 Se 在文件的开头会有{"pages": [,然后是所有缩进的项目行 {"title": "x", "text": "x", "tags": "x", "url": "x"},,最后是]}

【问题讨论】:

    标签: python json scrapy export scrapy-pipeline


    【解决方案1】:

    就内存使用而言,这不是一个好习惯,但一种选择是保留一个对象并在进程结束时将其写入:

    class RautahakuPipeline(object):
    
        def open_spider(self, spider):
            self.items = { "pages":[] }
            self.file = null # open('items.json', 'w')
    
        def close_spider(self, spider):
            self.file = open('items.json', 'w')
            self.file.write(json.dumps(self.items))
            self.file.close()
    
        def process_item(self, item, spider):            
            self.items["pages"].append(dict(item))
            return item
    

    然后,如果内存是一个问题(无论如何都必须注意),请尝试如下编写json文件:

    class RautahakuPipeline(object):
    
        def open_spider(self, spider):
            self.file = open('items.json', 'w')
            header='{"pages": ['
            self.file.write(header)
    
        def close_spider(self, spider):
            footer=']}'
            self.file.write(footer)
            self.file.close()
    
        def process_item(self, item, spider):
            line = json.dumps(dict(item)) + "\n"
            self.file.write(line)
            return item
    

    希望对你有帮助。

    【讨论】:

    • 谢谢!您的第一个建议非常有效!我唯一要挑剔的是每个项目都在同一行。我应该在哪里添加 \n 还是用其他方式完成?
    • 因为json.dumps 是一次写入的,所以您不能写入\n。在这种情况下,请尝试使用选项二。
    【解决方案2】:

    使用list comprehension。我不知道您的数据看起来如何,但使用一个玩具示例:

    product_title = range(1,10)
    product_text = range(10,20)
    product_tags = range(20,30)
    product_url = range(30,40)
    
    item = {'pages':[{'title':title,'text':text,'tags':tags,'url':url} 
      for title, text, tags, url in zip(product_title, product_text, product_tags, product_url)]}
    

    我得到这个结果:

    {'pages': [{'tags': 20, 'text': 10, 'title': 1, 'url': 30},
    {'tags': 21, 'text': 11, 'title': 2, 'url': 31},
    {'tags': 22, 'text': 12, 'title': 3, 'url': 32},
    {'tags': 23, 'text': 13, 'title': 4, 'url': 33},
    {'tags': 24, 'text': 14, 'title': 5, 'url': 34},
    {'tags': 25, 'text': 15, 'title': 6, 'url': 35},
    {'tags': 26, 'text': 16, 'title': 7, 'url': 36},
    {'tags': 27, 'text': 17, 'title': 8, 'url': 37},
    {'tags': 28, 'text': 18, 'title': 9, 'url': 38}]}
    

    【讨论】:

    • 我分享了我的蜘蛛的整个代码,因为它的结构可能是你的解决方案不适合我的原因。你能看看我的第二次编辑吗?
    【解决方案3】:
    items = {}
    #item = TechbbsItem()  # not sure what this is doing?
    items['pages'] = []
    for title, text, tags, url in zip(product_title, product_text, product_tags, product_url):
        temp_dict = {}
        temp_dict['title'] = title
        temp_dict['text'] = text
        temp_dict['tags'] = tags
        temp_dict['url'] = url
        items["pages"].append(temp_dict)
    return items
    

    【讨论】:

    • 运行蜘蛛会报错'AttributeError: dict' object has no attribute 'append'
    • @JanneSalmi 已修复
    • 我分享了我的蜘蛛的整个代码,因为它的结构可能是你的解决方案不适合我的原因。你能看看我的第二次编辑吗?
    猜你喜欢
    • 2016-07-25
    • 1970-01-01
    • 1970-01-01
    • 2011-12-11
    • 2015-02-27
    • 2015-10-06
    • 1970-01-01
    • 2014-02-19
    • 2017-08-18
    相关资源
    最近更新 更多