【问题标题】:Scrapy Crawls only 1st page and not the restScrapy 只抓取第一页而不是其余的
【发布时间】:2013-07-14 18:25:33
【问题描述】:

嘿,我正在使用 scrapy 制作一个项目,其中我需要从业务目录中删除业务详细信息 http://directory.thesun.co.uk/find/uk/computer-repair
我面临的问题是:当我尝试抓取页面时,我的爬虫仅获取第一页的详细信息,而我还需要获取其余 9 页的详细信息;那是所有10页.. 我在我的蜘蛛代码和 items.py 和设置 .py 下面显示 请查看我的代码并帮助我解决它

蜘蛛代码::

from scrapy.spider import BaseSpider
from scrapy.selector import HtmlXPathSelector
from project2.items import Project2Item

class ProjectSpider(BaseSpider):
    name = "project2spider"
    allowed_domains = ["http://directory.thesun.co.uk/"]
    start_urls = [
        "http://directory.thesun.co.uk/find/uk/computer-repair"
    ]

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        sites = hxs.select('//div[@class="abTbl "]')
        items = []
        for site in sites:
            item = Project2Item()
            item['Catogory'] = site.select('span[@class="icListBusType"]/text()').extract()
            item['Bussiness_name'] = site.select('a/@title').extract()
            item['Description'] = site.select('span[last()]/text()').extract()
            item['Number'] = site.select('span[@class="searchInfoLabel"]/span/@id').extract()
            item['Web_url'] = site.select('span[@class="searchInfoLabel"]/a/@href').extract()
            item['adress_name'] = site.select('span[@class="searchInfoLabel"]/span/text()').extract()
            item['Photo_name'] = site.select('img/@alt').extract()
            item['Photo_path'] = site.select('img/@src').extract()
            items.append(item)
        return items

我的 items.py 代码如下::

from scrapy.item import Item, Field

class Project2Item(Item):
    Catogory = Field()
    Bussiness_name = Field()
    Description = Field()
    Number = Field()
    Web_url = Field()
    adress_name = Field()
    Photo_name = Field()
    Photo_path = Field()

我的 settings.py 是:::

BOT_NAME = 'project2'

SPIDER_MODULES = ['project2.spiders']
NEWSPIDER_MODULE = 'project2.spiders'

请帮忙 我也从其他页面提取详细信息...

【问题讨论】:

    标签: python django scrapy


    【解决方案1】:

    以下是工作代码。滚动页面应该通过研究 网站及其滚动结构并相应地应用它们。在这种情况下,网站给了它“/page/x”,其中 x 是页码。

    from scrapy.spider import BaseSpider
    from scrapy.selector import HtmlXPathSelector
    from project2spider.items import Project2Item
    from scrapy.http import Request
    
    class ProjectSpider(BaseSpider):
        name = "project2spider"
        allowed_domains = ["http://directory.thesun.co.uk"]
        current_page_no = 1 
        start_urls = [ 
            "http://directory.thesun.co.uk/find/uk/computer-repair"
        ]   
    
        def get_next_url(self, fired_url):
            if '/page/' in fired_url:
                url, page_no = fired_url.rsplit('/page/', 1)
            else:
                if self.current_page_no != 1:
                    #end of scroll
                    return 
            self.current_page_no += 1
            return "http://directory.thesun.co.uk/find/uk/computer-repair/page/%s" % self.current_page_no
    
        def parse(self, response):
            fired_url = response.url
            hxs = HtmlXPathSelector(response)
            sites = hxs.select('//div[@class="abTbl "]')
            for site in sites:
                item = Project2Item()
                item['Catogory'] = site.select('span[@class="icListBusType"]/text()').extract()
                item['Bussiness_name'] = site.select('a/@title').extract()
                item['Description'] = site.select('span[last()]/text()').extract()
                item['Number'] = site.select('span[@class="searchInfoLabel"]/span/@id').extract()
                item['Web_url'] = site.select('span[@class="searchInfoLabel"]/a/@href').extract()
                item['adress_name'] = site.select('span[@class="searchInfoLabel"]/span/text()').extract()
                item['Photo_name'] = site.select('img/@alt').extract()
                item['Photo_path'] = site.select('img/@src').extract()
                yield item
            next_url = self.get_next_url(fired_url)
            if next_url:
                yield Request(next_url, self.parse, dont_filter=True)
    `
    

    【讨论】:

      【解决方案2】:

      如果您检查分页链接,它们看起来像这样:

      http://directory.thesun.co.uk/find/uk/computer-repair/page/3 http://directory.thesun.co.uk/find/uk/computer-repair/page/2

      您可以使用带有变量的 urllib2 循环页面

      import urllib2
      response = urllib2.urlopen('http://directory.thesun.co.uk/find/uk/computer-repair/page/' + page)
      html = response.read()
      

      然后抓取 html。

      【讨论】:

        【解决方案3】:

        我尝试了@nizam.sp 的代码。已发布,这仅显示来自主页的 2 条记录 1 条记录(最后一条记录)和来自第二页的 1 条记录(随机记录)并结束。

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2023-03-30
          • 2023-01-24
          • 1970-01-01
          • 1970-01-01
          • 2013-11-30
          • 2023-03-30
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多