【问题标题】:Scrapy spider returns different values compared to browser console xpath result与浏览器控制台 xpath 结果相比,Scrapy spider 返回不同的值
【发布时间】:2017-03-13 23:06:14
【问题描述】:

xpath:

//ol[@class="breadcrumb container"]/li[not(contains(@class,"first")) and not(contains(@class,"last"))]/a/span/text()

HTML:

<ol class="breadcrumb container">
    <li class="first"><a href="http://example.com/index.php?route=common/home"><span>Home</span></a></li>
    <li><a href="http://example.com/books"><span>Books</span></a></li>
    <li class="last"><a href="http://example.com/books?product_id=193" class="last"><span>My Vision : Challenges in the Race for Excellence - Mohammed Bin Rashid Al Maktoum</span></a></li>
</ol>

Python 代码:

categories = ['NO DATA', 'NO DATA', 'NO DATA', 'NO DATA', 'NO DATA', 'NO DATA']
catIndex = 0
for cat in sel.xpath('//ol[@class="breadcrumb container"]/li[not(contains(@class,"first")) and not(contains(@class,"last"))]/a/span/text()').extract():
            categories[catIndex] = cat
            catIndex += 1

想要的结果是“书籍”,当我使用 xpath 在 Firebug 控制台上检查它时,它会返回正确的结果,但是当我运行蜘蛛时,它会返回整个 3 Li 元素,不包括 class="first" 和 class="last"

我尝试了命令 Scrapy View http://example.com 来查看页面蜘蛛如何看到它,但一切看起来都一样,xpath 返回正确的结果

当我尝试在 Scrapy Shell 中使用 xpath 时,它返回的所有 3 个 Li 元素的结果都不正确

可能是什么问题?

【问题讨论】:

  • 实际上,如果您从浏览器检查页面源,它​​会告诉您第一个和最后一个 li 元素没有类属性,因此最终您将在结果中获得所有三个元素。这就是问题所在。
  • 您需要在您的 xpath 中进行更改以获得您想要的正确输出。
  • 你说得对
  • 没有Class属性

标签: python xpath scrapy


【解决方案1】:

在 Internet Explorer 中打开 Scrapy View http://example.com 输出,发现 Li 元素中没有 Class 属性。

似乎在 Chrome 或 Firefox 中打开的 Scrapy View 命令没有显示蜘蛛看到的真实代码。

【讨论】:

    猜你喜欢
    相关资源
    最近更新 更多
    热门标签