【发布时间】:2013-08-07 03:42:55
【问题描述】:
以下用于抓取 vBulletin 论坛网站的简单代码遇到问题:
class ForumSpider(CrawlSpider):
...
rules = (
Rule(SgmlLinkExtractor(restrict_xpaths="//div[@class='threadlink condensed']"),
callback='parse_threads'),
)
def parse_threads(self, response):
thread = HtmlXPathSelector(response)
# get the list of posts
posts = thread.select("//div[@id='posts']//table[contains(@id,'post')]/*")
# plist = []
for p in posts:
table = ThreadItem()
table['thread_id'] = (p.select("//input[@name='searchthreadid']/@value").extract())[0].strip()
string_id = p.select("../@id").extract() # returns a list
p_id = string_id[0].split("post")
table['post_id'] = p_id[1]
# plist.append(table)
# return plist
yield table
除了一些 xpath hackiness 之外,当我使用 yield 运行它时,我会得到非常奇怪的结果,多次点击相同的 thread_id 和 post_id。比如:
114763,1314728
114763,1314728
114763,1314728
114763,1314740
114763,1314740
114763,1314740
当我使用 return (在 cmets 中)切换回相同的逻辑时,一切正常。我认为这可能是生成器的一些基本错误,但我无法弄清楚。为什么相同的帖子会被多次点击?为什么代码使用 return 而不是 yield 工作?
gist here 中的完整代码 sn-p。
【问题讨论】: