【发布时间】:2016-02-26 00:52:00
【问题描述】:
场景:
- 具有多个蜘蛛的单个 scrapy 项目。
- 所有蜘蛛从脚本一起运行。
问题:
- 同一命名空间中的所有日志消息。不可能知道哪条消息属于哪个蜘蛛。
在 scrapy 0.24 中,我有多个蜘蛛在一个脚本中运行,我得到一个日志文件,其中包含与其蜘蛛相关的消息,类似于:
2015-09-30 22:55:12-0400 [scrapy] INFO: Scrapy 0.24.5 started (bot: mybot)
2015-09-30 22:55:12-0400 [scrapy] DEBUG: Enabled extensions: LogStats, ...
2015-09-30 21:55:12-0500 [scrapy] DEBUG: Enabled downloader middlewares: HttpAuthMiddleware, ...
2015-09-30 21:55:12-0500 [scrapy] DEBUG: Enabled spider middlewares: HttpErrorMiddleware, ...
2015-09-30 21:55:12-0500 [scrapy] DEBUG: Enabled item pipelines: MybotPipeline
2015-09-30 21:55:12-0500 [spider1] INFO: Spider opened
2015-09-30 21:55:12-0500 [spider1] INFO: Crawled 0 pages ...
2015-09-30 21:55:12-0500 [spider2] INFO: Spider opened
2015-09-30 21:55:12-0500 [spider2] INFO: Crawled 0 pages ...
2015-09-30 21:55:12-0500 [spider3] INFO: Spider opened
2015-09-30 21:55:12-0500 [spider3] INFO: Crawled 0 pages ...
2015-09-30 21:55:13-0500 [spider2] DEBUG: Crawled (200) <GET ...
2015-09-30 21:55:13-0500 [spider3] DEBUG: Crawled (200) <GET ...
2015-09-30 21:55:13-0500 [spider1] DEBUG: Crawled (200) <GET ...
2015-09-30 21:55:13-0500 [spider1] INFO: Closing spider (finished)
2015-09-30 21:55:13-0500 [spider1] INFO: Dumping Scrapy stats: ...
2015-09-30 21:55:13-0500 [spider3 INFO: Closing spider (finished)
2015-09-30 21:55:13-0500 [spider3] INFO: Dumping Scrapy stats: ...
2015-09-30 21:55:13-0500 [spider2] INFO: Closing spider (finished)
2015-09-30 21:55:13-0500 [spider2] INFO: Dumping Scrapy stats: ...
有了这个日志文件,我可以在需要时运行grep spiderX logfile.txt 来获取与特定蜘蛛相关的日志。但现在,在 scrapy 1.0 中,我得到了:
2015-09-30 21:55:12-0500 [scrapy] INFO: Spider opened
2015-09-30 21:55:12-0500 [scrapy] INFO: Crawled 0 pages ...
2015-09-30 21:55:12-0500 [scrapy] INFO: Spider opened
2015-09-30 21:55:12-0500 [scrapy] INFO: Crawled 0 pages ...
2015-09-30 21:55:12-0500 [scrapy] INFO: Spider opened
2015-09-30 21:55:12-0500 [scrapy] INFO: Crawled 0 pages ...
显然不可能知道每个蜘蛛属于哪条消息。
问题是:有没有办法让以前的行为?
每个蜘蛛也可以有不同的日志文件。 [1]
但是使用custom_settings 覆盖蜘蛛中的日志文件是不可能的。 [2]
那么,有没有办法为每个蜘蛛创建不同的日志文件?
[1]Scrapy Project with Multiple Spiders - Custom Settings Ignored
[2]https://github.com/scrapy/scrapy/issues/1612
【问题讨论】: