【问题标题】:I want to read from a file and scrape each URL我想从文件中读取并抓取每个 URL
【发布时间】:2019-02-07 08:55:08
【问题描述】:

我想要做的是从文件中读取每个 URL 并抓取该 URL。之后,我将抓取数据移动到类WebRealTor,然后将数据序列化为json,最后将所有数据保存在一个json文件中。 这是文件的内容: https://www.seloger.com/annonces/achat/appartement/paris-14eme-75/montsouris-dareau/143580615.htm?ci=750114&idtt=2,5&idtypebien=2,1&LISTING-LISTpg=8&naturebien=1,2,4&tri=initial&bd=ListToDetail https://www.seloger.com/annonces/achat/appartement/montpellier-34/gambetta/137987697.htm?ci=340172&idtt=2,5&idtypebien=1,2&naturebien=1,2,4&tri=initial&bd=ListToDetail https://www.seloger.com/annonces/achat/appartement/montpellier-34/celleneuve/142626025.htm?ci=340172&idtt=2,5&idtypebien=1,2&naturebien=1,2,4&tri=initial&bd=ListToDetail https://www.seloger.com/annonces/achat/appartement/versailles-78/domaine-national-du-chateau/138291887.htm

我的脚本是:

import scrapy
import json



class selogerSpider(scrapy.Spider):
    name = "realtor"

    custom_settings = {
        'DOWNLOADER_MIDDLEWARES': {
            'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None,
            'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware': 400,
    }
}

    def start_requests(self):
         with open("annonces.txt", "r") as file:
             for line in file.readlines():
                  yield scrapy.Request(line)
    def parse(self, response):
        name = response.css(".agence-link::text").extract_first()
        address = response.css(".agence-adresse::text").extract_first()

        XPATH_siren = ".//div[@class='legalNoticeAgency']//p/text()"
        siren = response.xpath(XPATH_siren).extract_first()

        XPATH_website = ".//div[@class='agence-links']//a/@href"
        site = response.xpath(XPATH_website).extract()

        XPATH_phone = ".//div[@class='contact g-row-50']//div[@class='g-col g-50 u-pad-0']//button[@class='btn-phone b-btn b-second fi fi-phone tagClick']/@data-phone"
        phone = response.xpath(XPATH_phone).extract_first()


        yield {
            'Agency_Name =': name,
            'Agency_Address =': address,
            'Agency_profile_website =': site,
            'Agency_number =': phone,
            'Agency_siren =': siren
        }

        file.close()


class WebRealTor:

    def __name__(self):
        self.nom = selogerSpider.name
    def __address__(self):
        self.adress = selogerSpider.address
    def __sirenn__(self):
        self.sire = selogerSpider.siren
    def __numero__(self):
        self.numero = selogerSpider.phone


with open('data.txt', 'w') as outfile:
    json.dump(data, outfile)

【问题讨论】:

  • 请注意,本网站不授权报废:)。但是,您的问题到底是什么?什么不起作用?
  • 好的,我希望它将抓取数据移动到 webRealTor 类,然后将其保存在 json 文件中
  • 打开文件 ---> 抓取数据 ---> 类 webrealtor ---> 序列化 json 中的数据 ----> 将数据保存到 json 文件中
  • 那么你的代码就差不多了,这有点简单。使用readlines 然后对于每一行,使用返回数据extract_data 的自定义方法,然后将它们添加到您的类中,然后创建一个字典,然后将其保存到一个 json 文件中>
  • @BlueSheepToken 我是 python 初学者的问题,所以我需要帮助编辑我的代码

标签: python scrapy


【解决方案1】:

尝试将所有内容移至课堂上的start_requests。像这样:

def start_requests(self):
    with open("annonces.txt", "r") as file:
        for line in file.readlines():
            yield scrapy.Request(line)  # self.parse is by default

def parse(self, response):
    # each link parsing as you already did

【讨论】:

  • 然后他将如何刮掉每一行?
  • 您的文件中有链接。这将对文件中的每个链接发出请求。你能详细说明你的问题吗?
  • 当我输入这段代码时,他用红线在 Request 下划线
  • 好的,然后我想将数据移动到 webrealtor 类,然后将其保存在 json 文件中
  • @joes 看看将 Scrapy 的 Item 用于您的数据类。要导出 JSON 文件,您可以使用 Feed exporter。这是一个简单的export example
猜你喜欢
  • 1970-01-01
  • 2019-03-27
  • 2021-01-13
  • 1970-01-01
  • 1970-01-01
  • 2018-06-23
  • 2021-09-25
  • 1970-01-01
  • 2010-09-09
相关资源
最近更新 更多