【发布时间】:2016-10-30 20:06:15
【问题描述】:
我想在这个网站上提取数据:http://www.pokepedia.fr/Pikachu 我正在学习 python 以及如何使用 Scrapy,我的问题是:为什么我不能用 Xpath 检索数据?
当我在浏览器中测试 xpath 时,我的 Xpath 看起来不错,它返回了正确的值。 (谷歌浏览器)
import re
from scrapy import Spider
from scrapy.selector import Selector
from stack.items import StackItem
class StackSpider(Spider):
name = "stack"
allowed_domains = ["pokepedia.fr"]
start_urls = [
"http://www.pokepedia.fr/Pikachu",
]
def unicodize(seg):
if re.match(r'\\u[0-9a-f]{4}', seg):
return seg.decode('unicode-escape')
return seg.decode('utf-8')
def parse(self, response):
pokemon = Selector(response).xpath('//*[@id="mw-content-text"]/table[2]')
for question in pokemon:
item = StackItem()
item['title'] = question.xpath(
'//*[@id="mw-content-text"]/table[2]/tbody/tr[1]/th[2]/text()').extract()[0]
yield item
我想在页面中提取口袋妖怪的名称,但是当我使用时:
scrapy crawl stack -o items.json -t json
我的 Json 输出:
[
在我的控制台中我有这个错误:
IndexError : list index out of range
我已经关注了这个教程:https://realpython.com/blog/python/web-scraping-with-scrapy-and-mongodb/
【问题讨论】:
-
如提供的答案中所述,请小心信任任何 Web 浏览器开发控制台/xpath 查看器,因为它们显示的文档并不总是页面生成的确切 HTML。通常它会添加标签,并修复任何损坏的 html。通常最好直接下载页面的 html(简单的 python 脚本可以做到)并从中获取单词。 Web Scraping 是一个很好的学习工具,但请始终牢记这个提示,它已经让我痛苦了好几次。
标签: python xpath web-scraping scrapy