【问题标题】:How to use multiple proxy when crawling with scrapy + splash?使用scrapy + splash爬取时如何使用多个代理?
【发布时间】:2016-07-14 04:54:25
【问题描述】:

我们使用 scrapy + splash 进行爬网,并且我们希望使用多个代理。但是splash只支持单代理https://splash.readthedocs.io/en/stable/api.html#proxy-profiles。

[proxy]

; required
host=proxy.crawlera.com
port=8010

; optional, default is no auth
username=username
password=password

; optional, default is HTTP. Allowed values are HTTP and SOCKS5
type=HTTP

scrapy+splash爬取时如何使用多个代理?

【问题讨论】:

  • “只支持一个代理”是什么意思?这不是真的,您可以使用不同的代理创建各种配置文件,并且对于每个请求,您可以告诉爬虫您要使用哪个配置文件,使爬虫使用不同的进程

标签: proxy scrapy scrapy-splash


【解决方案1】:

有几种选择:

  • 使用多个配置文件(正如 Rafael Almeida 在评论中建议的那样);
  • 为每个请求传递不同的代理 URL(请参阅 http://splash.readthedocs.io/en/stable/api.html#arg-proxy);
  • 编写一个 Splash Lua 脚本并在 splash:on_request 回调中使用 request:set_proxy - 文档中有一个示例。这样,您可以为页面初始化的不同请求设置不同的代理,而不仅仅是每个呈现页面的单个代理。我不知道在其他浏览器自动化工具(如 phantomjs 或 selenium)中可以做到这一点。

【讨论】:

    猜你喜欢
    • 2019-05-28
    • 1970-01-01
    • 2017-09-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-06-30
    • 2021-12-26
    • 2014-08-22
    相关资源
    最近更新 更多