【发布时间】:2018-02-23 13:58:00
【问题描述】:
我发现我的问题和scrapyd deploy shows 0 spiders很相似。我也多次尝试接受的答案,但它对我不起作用,所以我来寻求帮助。
项目目录为timediff_crawler,目录树形图为:
timediff_crawler/
├── scrapy.cfg
├── scrapyd-deploy
├── timediff_crawler
│ ├── __init__.py
│ ├── items.py
│ ├── pipelines.py
│ ├── settings.py
│ ├── spiders
│ │ ├── __init__.py
│ │ ├── prod
│ │ │ ├── __init__.py
│ │ │ ├── job
│ │ │ │ ├── __init__.py
│ │ │ │ ├── zhuopin.py
│ │ │ ├── rent
│ │ │ │ ├── australia_rent.py
│ │ │ │ ├── canada_rent.py
│ │ │ │ ├── germany_rent.py
│ │ │ │ ├── __init__.py
│ │ │ │ ├── korea_rent.py
│ │ │ │ ├── singapore_rent.py
...
1.1启动scrapyd,就ok了
(crawl_env)web@ha-2:/opt/crawler$ scrapyd
2015-11-11 15:00:37+0800 [-] Log opened.
2015-11-11 15:00:37+0800 [-] twistd 15.4.0 (/opt/crawler/crawl_env/bin/python 2.7.6) starting up.
2015-11-11 15:00:37+0800 [-] reactor class: twisted.internet.epollreactor.EPollReactor.
2015-11-11 15:00:37+0800 [-] Site starting on 6800
...
1.2 编辑scrapy.cfg
[settings]
default = timediff_crawler.settings
[deploy:ha2-crawl]
url = http://localhost:6800/
project = timediff_crawler
1.3 部署项目
(crawl_env)web@ha-2:/opt/crawler/timediff_crawler$ ./scrapyd-deploy -l
ha2-crawl http://localhost:6800/
(crawl_env)web@ha-2:/opt/crawler/timediff_crawler$ ./scrapyd-deploy ha2-crawl -p timediff_crawler
Packing version 1447229952
Deploying to project "timediff_crawler" in http://localhost:6800/addversion.json
Server response (200):
{"status": "ok", "project": "timediff_crawler", "version": "1447229952", "spiders": 0, "node_name": "ha-2"}
1.4 问题
响应显示蜘蛛的数量是0,实际上我有大约10只蜘蛛。
我按照这篇帖子scrapyd deploy shows 0 spiders中的建议,删除所有项目、版本、目录(包括build/egg/project.egg-info setup.py)并尝试再次部署,但它不起作用,蜘蛛的数量始终为 0。
我验证了 egg 文件,输出显示它似乎没问题:
(crawl_env)web@ha-2:/opt/crawler/timediff_crawler/eggs/timediff_crawler$ unzip -t 1447229952.egg
Archive: 1447229952.egg
testing: timediff_crawler/pipelines.py OK
testing: timediff_crawler/__init__.py OK
testing: timediff_crawler/items.py OK
testing: timediff_crawler/spiders/prod/job/zhuopin.py OK
testing: timediff_crawler/spiders/prod/rent/singapore_rent.py OK
testing: timediff_crawler/spiders/prod/rent/australia_rent.py OK
...
所以我不知道出了什么问题,请帮助并提前感谢!
【问题讨论】:
-
你能分享你的设置文件吗?创建一个pastebin
-
@eLRuLL,这里:settings.py
-
尝试将
SPIDER_MODULES更改为蜘蛛所在的路径:例如['timediff_crawler.spiders.prod.job']。 -
@eLRuLL 我又试了一次,还是不行。我也试过把所有spider源文件直接放到
timediff_crawler/spiders里,也不行。 -
当你运行
scrapy list时,它会列出你的蜘蛛吗?
标签: scrapy scrapy-spider