【问题标题】:Scrapy generate csv file (UTF-8)Scrapy 生成 csv 文件 (UTF-8)
【发布时间】:2017-05-01 10:49:16
【问题描述】:

我尝试使用爬虫的结果生成一个 CSV 文件。因为它是德语,所以我需要对其进行 UTF-8 编码(ä、ö 等)。这是我到目前为止的结果:

蜘蛛.py

import scrapy

from scrapy.spiders import BaseSpider
from scrapy.selector import Selector
from Polizeimeldungen.items import PolizeimeldungenItem


class PoliceSpider(scrapy.Spider):
  name = "pm"
  allowed_domains = ["berlin.de"]
  start_urls = 
["https://www.berlin.de/polizei/polizeimeldungen/archiv/2014/?page_at_1_0=1"]

  def parse(self, response):
    for sel in response.css('.row-fluid'):
        item = PolizeimeldungenItem()
        item['title'] = sel.css('a ::text').extract_first().encode('utf-8')
        item['link'] = sel.css('a ::text').extract_first().encode('utf-8') // this is wrong, but it is easy to fix  
        yield item

items.py

import scrapy

class PolizeimeldungenItem(scrapy.Item):
    title = scrapy.Field()
    link = scrapy.Field()

管道.py

import csv
class PolizeimeldungenPipeline(object):
def __init__(self):
    self.myCsv = csv.writer(open('Item.csv', 'wb'))
    self.myCsv.writerow(['title', 'link'])

    def process_item(self, item, spider):          
        self.myCsv.writerow([item['title'], item['link']])
        return item

Settings.py

BOT_NAME = 'Polizeimeldungen'

SPIDER_MODULES = ['Polizeimeldungen.spiders']
NEWSPIDER_MODULE = 'Polizeimeldungen.spiders'
ITEM_PIPELINES = {'Polizeimeldungen.pipelines.PolizeimeldungenPipeline': 100}

作为之后的结果:

scrapy crawl pm

我收到此错误消息:

TypeError: a bytes-like object is required, not 'str'

感谢您的帮助!!

更新:Python 3.6.0 :: Anaconda 4.3.1

【问题讨论】:

  • 您为什么不使用 scrapy crawl pm -o output_file.csv 来使用内置 csv 序列化的任何具体原因?
  • 另外,您自定义构建的输出管道中使用的 csv lib 似乎需要一些调整才能进行 utf-8 输出:docs.python.org/2/library/csv.html#examples
  • @rrschmidt 这仅适用于 Python 2。显然(这是一件好事),OP 正在使用 Python 3,否则 TypeError 消息将毫无意义(在 Python 2 中 str 是字节类对象)。

标签: python csv utf-8 scrapy


【解决方案1】:

我假设您使用的是 Python 3(此解决方案不适用于 Python 2)。

你需要改变两件事:

  • 使用所需的输出编码以文本模式打开输出文件。 在PolizeimeldungenPipeline的构造函数中,写:

    self.myCsv = csv.writer(open('Item.csv', 'w', encoding='utf-8'))
    
  • 不要对单元格进行编码(如PoliceSpider.parse):

    item['title'] = sel.css('a ::text').extract_first()
    

    等等

【讨论】:

  • 谢谢@lenz 有效积分!它可以正常工作,但 CSV 文件仍然是空的 :(
  • 到目前为止,您还没有提到输出为空。这很可能是一个不同的问题。
猜你喜欢
  • 2018-01-16
  • 2021-10-10
  • 1970-01-01
  • 2019-11-28
  • 2023-03-15
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多