【问题标题】:Scrapy a website that need use cookies抓取一个需要使用 cookie 的网站
【发布时间】:2014-04-26 06:26:15
【问题描述】:

我正在制作用于抓取网站的scrapy,但该网站正在使用cookies,我不知道如何使用cookies制作用于抓取网站数据的指令

class DmozSpider(Spider):
    name = "dmoz"
    allowed_domains = ["dmoz.org"]
    start_urls = [
        "http://www.dmoz.org/Computers/Programming/Languages/Python/Books/"
    ]

    def parse(self, response):
        sel = Selector(response)
        sites = sel.HtmlXPathSelector('//ul[@class="directory-url"]/li')
        items = []

        for site in sites:
            item = Website()
            item['name'] = site.xpath('a/text()').extract()
            item['url'] = site.xpath('a/@href').extract()
            items.append(item)

        return items

如何将 cookie 正确添加到此网址

【问题讨论】:

    标签: python cookies xpath web-scraping scrapy


    【解决方案1】:

    要扩展@omair_77 的答案,您可以覆盖蜘蛛的start_requests 方法,将cookie 添加到蜘蛛的初始请求中:

    def start_requests(self):
        return [Request(url="http://www.example.com",
               cookies={'currency': 'USD', 'country': 'UY'})]
    

    这样,您的蜘蛛将使用这些 cookie 发出第一个请求,而对您的 parse 方法的第一次调用将带有响应。

    http://scrapy.readthedocs.org/en/latest/topics/spiders.html#scrapy.spider.Spider.start_requests

    【讨论】:

      【解决方案2】:

      你可以像这样添加cookie

      request_with_cookies = Request(url="http://www.example.com",
                     cookies={'currency': 'USD', 'country': 'UY'})
      

      http://doc.scrapy.org/en/latest/topics/request-response.html#topics-request-response

      【讨论】:

      • 谢谢,我试试这个,但我不能迭代这个 request_with_cookies
      • 您能否详细说明一下,您所说的迭代是什么意思?
      • 有了那个 request_wit_cookies,我无法在使用 xpath 获取数据后获取项目
      猜你喜欢
      • 2011-07-03
      • 1970-01-01
      • 2021-04-22
      • 1970-01-01
      • 2021-12-31
      • 1970-01-01
      • 1970-01-01
      • 2021-04-17
      相关资源
      最近更新 更多