【问题标题】:Newbie: How to scrape multiple web pages with only one start_urls?新手:如何只用一个 start_urls 抓取多个网页?
【发布时间】:2013-06-03 15:09:25
【问题描述】:

首先,我试图从以下位置获取资金代码,例如 MGB_U、JAS_U: "http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd=MMFU_U"

然后,从示例中提取每个基金的价格:

"http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd="+"MGB_U"

"http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd="+"JAS_U"

我的代码有: raise NotImplementedError 但我还是不知道怎么解决。

from scrapy.spider import BaseSpider
from scrapy.selector import HtmlXPathSelector
from fundPrice.items import FundPriceItem

class PruSpider(BaseSpider):
    name = "prufunds"
    allowed_domains = ["prudential.com.hk"]
    start_urls = ["http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd=MMFU_U"]

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        funds_U = hxs.select('//table//table//table//table//select[@class="fundDropdown"]//option//@value').extract()
        funds_U = [x for x in funds_U if x != (u"#" and u"MMFU_U")]

        items = []

        for fund_U in funds_U:
            url = "http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd=" + fund_U
            item = FundPriceItem()
            item['fund'] = fund_U
            item['data'] =  hxs.select('//table//table//table//table//td[@class="fundPriceCell1" or @class="fundPriceCell2"]//text()').extract()
            items.append(item)
            return items

【问题讨论】:

    标签: python web-scraping scrapy


    【解决方案1】:

    您应该为循环中的每个fund 使用scrapy 的Request:

    from scrapy.http import Request
    from scrapy.spider import BaseSpider
    from scrapy.selector import HtmlXPathSelector
    from fundPrice.items import FundPriceItem
    
    
    class PruSpider(BaseSpider):
        name = "prufunds"
        allowed_domains = ["prudential.com.hk"]
        start_urls = ["http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd=MMFU_U"]
    
        def parse(self, response):
            hxs = HtmlXPathSelector(response)
            funds_U = hxs.select('//table//table//table//table//select[@class="fundDropdown"]//option//@value').extract()
            funds_U = [x for x in funds_U if x != (u"#" and u"MMFU_U")]
    
            for fund_U in funds_U:
                yield Request(
                    url="http://www.prudential.com.hk/PruServlet?module=fund&purpose=searchHistFund&fundCd=" + fund_U,
                    callback=self.parse_fund,
                    meta={'fund': fund_U})
    
        def parse_fund(self, response):
            hxs = HtmlXPathSelector(response)
            item = FundPriceItem()
            item['fund'] = response.meta['fund']
            item['data'] = hxs.select(
                '//table//table//table//table//td[@class="fundPriceCell1" or @class="fundPriceCell2"]//text()').extract()
            return item
    

    希望对您有所帮助。

    【讨论】:

    • 从 item['fund'] = response.meta['fund'],资金字符串将是例如MGB_U,JAS_U。如何删除 _U 并将其保留为 MGB、JAS 等?
    • 就像response.meta['fund'].split('_')[0]。如果有帮助,也请考虑接受答案。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-09-10
    • 1970-01-01
    • 2021-06-03
    • 2020-03-27
    • 2021-03-10
    相关资源
    最近更新 更多