【发布时间】:2015-07-02 17:33:08
【问题描述】:
我正在为 wunderground.com 构建一个网络爬虫,但我的代码返回了英寸雨量和湿度的“[]”值。谁能明白为什么会这样?
# -*- coding: utf-8 -*-
import scrapy
from scrapy.selector import Selector
import time
from wunderground_scraper.items import WundergroundScraperItem
class WundergroundComSpider(scrapy.Spider):
name = "wunderground"
allowed_domains = ["www.wunderground.com"]
start_urls = (
'http://www.wunderground.com/q/zmw:10001.5.99999',
)
def parse(self, response):
info_set = Selector(response).xpath('//div[@id="current"]')
list = []
for i in info_set:
item = WundergroundScraperItem()
item['description'] = i.xpath('div/div/div/div/span/text()').extract()
item['description'] = item['description'][0]
item['humidity'] = i.xpath('div/table/tbody/tr/td/span/span/text()').extract()
item['inches_rain'] = i.xpath('div/table/tbody/tr/td/span/span/text()').extract()
list.append(item)
return list
我也知道湿度和英寸雨量项设置为相同的 xpath,但这应该是正确的,因为一旦信息在数组中,我只需将它们设置为数组中的某些值。
【问题讨论】:
标签: python xpath web-scraping scrapy scrapy-spider