【发布时间】:2017-03-13 23:06:14
【问题描述】:
xpath:
//ol[@class="breadcrumb container"]/li[not(contains(@class,"first")) and not(contains(@class,"last"))]/a/span/text()
HTML:
<ol class="breadcrumb container">
<li class="first"><a href="http://example.com/index.php?route=common/home"><span>Home</span></a></li>
<li><a href="http://example.com/books"><span>Books</span></a></li>
<li class="last"><a href="http://example.com/books?product_id=193" class="last"><span>My Vision : Challenges in the Race for Excellence - Mohammed Bin Rashid Al Maktoum</span></a></li>
</ol>
Python 代码:
categories = ['NO DATA', 'NO DATA', 'NO DATA', 'NO DATA', 'NO DATA', 'NO DATA']
catIndex = 0
for cat in sel.xpath('//ol[@class="breadcrumb container"]/li[not(contains(@class,"first")) and not(contains(@class,"last"))]/a/span/text()').extract():
categories[catIndex] = cat
catIndex += 1
想要的结果是“书籍”,当我使用 xpath 在 Firebug 控制台上检查它时,它会返回正确的结果,但是当我运行蜘蛛时,它会返回整个 3 Li 元素,不包括 class="first" 和 class="last"
我尝试了命令 Scrapy View http://example.com 来查看页面蜘蛛如何看到它,但一切看起来都一样,xpath 返回正确的结果
当我尝试在 Scrapy Shell 中使用 xpath 时,它返回的所有 3 个 Li 元素的结果都不正确
可能是什么问题?
【问题讨论】:
-
实际上,如果您从浏览器检查页面源,它会告诉您第一个和最后一个 li 元素没有类属性,因此最终您将在结果中获得所有三个元素。这就是问题所在。
-
您需要在您的 xpath 中进行更改以获得您想要的正确输出。
-
你说得对
- 没有Class属性