【问题标题】:Recursive Scraping using scrapy python使用scrapy python进行递归抓取
【发布时间】:2018-01-14 05:24:03
【问题描述】:

我已经做了一个刮板,它可以从多个页面中刮取数据。 我的问题是我有一堆网址(比如大约 10 个网址),每次都需要传递。

这是我的代码,

# -*- coding: utf-8 -*-
import scrapy
import csv
import re
import sys
import os
from scrapy.linkextractor import LinkExtractor
from scrapy.spiders import Rule, CrawlSpider
from datablogger_scraper.items import DatabloggerScraperItem


class DatabloggerSpider(CrawlSpider):
    # The name of the spider
    name = "datablogger"

    # The domains that are allowed (links to other domains are skipped)
    allowed_domains = ["cityofalabaster.com"]
    print type(allowed_domains)

    # The URLs to start with
    start_urls = ["http://www.cityofalabaster.com/"]
    print type(start_urls)

    # This spider has one rule: extract all (unique and canonicalized) links, follow them and parse them using the parse_items method
    rules = [
        Rule(
            LinkExtractor(
                canonicalize=True,
                unique=True
            ),
            follow=True,
            callback="parse_items"
        )
    ]

    # Method which starts the requests by visiting all URLs specified in start_urls
    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url, callback=self.parse, dont_filter=True)

    # Method for parsing items
    def parse_items(self, response):
        # The list of items that are found on the particular page
        items = []
        # Only extract canonicalized and unique links (with respect to the current page)
        links = LinkExtractor(canonicalize=True, unique=True).extract_links(response)
        # Now go through all the found links
        for link in links:
            # Check whether the domain of the URL of the link is allowed; so whether it is in one of the allowed domains
            is_allowed = False
            for allowed_domain in self.allowed_domains:
                if allowed_domain in link.url:
                    is_allowed = True
            # If it is allowed, create a new item and add it to the list of found items
            if is_allowed:
                item = DatabloggerScraperItem()
                item['url_from'] = response.url
                item['url_to'] = link.url
                items.append(item)
        # Return all the found items
        return items

如果你看看我的代码, 您可以看到允许的域和 start_urls “链接”是手动传递的。 相反,我有包含要传递的 url 的 csv。

输入:-

http://www.daphneal.com/
http://www.digitaldecatur.com/
http://www.demopolisal.com/
http://www.dothan.org/
http://www.cofairhope.com/
http://www.florenceal.org/
http://www.fortpayne.org/
http://www.cityofgadsden.com/
http://www.cityofgardendale.com/
http://cityofgeorgiana.com/Home/
http://www.goodwater.org/
http://www.guinal.org/
http://www.gulfshoresal.gov/
http://www.guntersvilleal.org/index.php
http://www.hartselle.org/
http://www.headlandalabama.org/
http://www.cityofheflin.org/
http://www.hooveral.org/

这是将 url 和域传递给 Start_urls 和 allowed_domains 的代码。

import csv
import re
import sys
import os

with open("urls.csv") as csvfile:
        csvreader = csv.reader(csvfile, delimiter=",")
    for line in csvreader:
        start_urls = line[0]
        start_urls1 = start_urls.split()
        print start_urls1
        print type(start_urls1)
    if start_urls[7:10] == 'www':
        p = re.compile(ur'(?<=http://www.).*(?=\/|.*)')

    elif start_urls[7:10] != 'www' and start_urls[-1] == '/' :
        p = re.compile(ur'(?<=http://).*(?=\/|\s)')

    elif start_urls[7:10] != 'www' and start_urls[-1] != '/' :
        p = re.compile(ur'(?<=http://).*(?=\/|.*)')
    else:
        p = re.compile(ur'(?<=https://).*(?=\/|.*)')


        allowed_domains = re.search(p,start_urls).group()
        allowed_domains1 = allowed_domains.split()
        print allowed_domains1
        print type(allowed_domains1)

上面的代码将读取每个 url ,将每个 url 转换为列表(格式)并传递给 start_url 并通过应用正则表达式获取域并将其传递给 allowed_domain (格式)

我应该如何将上述代码集成到我的主代码中以避免手动传递 allowed_domains 和 start_urls ???

提前致谢!!!!

【问题讨论】:

    标签: python scrapy


    【解决方案1】:

    您可以从 python 脚本运行蜘蛛,查看更多 here:

    if __name__ == '__main__':
        from scrapy.crawler import CrawlerProcess
        process = CrawlerProcess({
            'USER_AGENT': 'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1)'
        })
    
        # parse from csv file
        allowed_domains = ...
        start_urls = ...
    
        DatabloggerSpider.allowed_domains = allowed_domains
        DatabloggerSpider.start_urls = start_urls
        process.crawl(DatabloggerSpider)
        process.start()
    

    【讨论】:

    • 所以你的意思是说,这是一个单独的python文件,用于读取csv url并调用scrapy主程序???我说的对吗?
    • 可以将解析逻辑包装到函数中,在这里调用。您获取这些值的方式可能会有所不同,这只是从 python 脚本运行蜘蛛的一种方式,而不是使用强制您手动设置这些东西的 scrapy shell。
    猜你喜欢
    • 1970-01-01
    • 2014-04-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-09-15
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多