【发布时间】:2014-05-21 20:34:16
【问题描述】:
我正在尝试从此网页 (http://www.sacbee.com/statepay/#req=employee%2Fsearch%2Fname%3D%2Fyear%3D2013%2Fdepartment%3DCSU%20Sacramento) 中提取 csu 员工工资数据。我尝试过使用 urlib2 和 requests 库,但它们都没有从网页返回实际的表格。我猜原因可能是该表是由 javascript 动态生成的。下面是我使用请求的代码。
from lxml import html
import requests
page = requests.get("http://www.sacbee.com/statepay/#req=employee%2Fsearch%2Fname%3D%2Fyear%3D2013%2Fdepartment%3DCSU%20Sacramento")
tree = html.fromstring(page.text)
name = tree.xpath('//table/tbody/tr/td[2]/text()'
任何帮助/cmets 将不胜感激。
【问题讨论】:
-
如果您检查页面,信息实际上是一个 JSON 文件。我希望你知道这意味着什么。 :D
-
谢谢七无!我知道如何处理 json 文件,但你能指出我的 json 文件吗?我无法在网页中找到 json 文件的 url。
-
网址是
http://api.sacbeelabs.com/v1/statepay/employee/search/name=/year=2013/department=CSU%20Sacramento.json。但是,您需要为此发出 POST 请求,因为它只会在 Python 中返回以下内容:{u'status': {u'message': u'Unauthorized', u'code': 401, u'reason': u'Client origin not specified'}, u'request': {u'verb': u'statepay/employee/search/name=/year=2013/department=CSU%20Sacramento', u'params': [], u'format': u'json'}}。 -
嗨 Nanashi,你是如何找到 json 文件的?
标签: python html web-crawler lxml scrape