【发布时间】:2016-05-02 10:00:42
【问题描述】:
我一直在研究一个网站:
http://directory.ccnecommunity.org/reports/rptAccreditedPrograms_New.asp?sort=institution
我需要从各个大学提取硕士下的数据。
您可能注意到并非每所大学都有硕士数据,所以我需要跟踪它。
在这种情况下如何跟踪数据?
到目前为止,我的 python 和 XPATH 代码:
import __future__
from lxml import html
import requests
from bs4 import BeautifulSoup
page = requests.get('http://directory.ccnecommunity.org/reports/rptAccreditedPrograms_New.asp?sort=institution')
soup = str(BeautifulSoup(page.content, 'html.parser'))
tree = html.fromstring(soup)
for table in tree.xpath('//table[@width="95%" and @align="center" and @class="center"]'):
print('-- NEW TABLE -- \n')
tab = table.xpath('.//table[@width="260px"]/tr/td[@style="width: 100%;"]/text()')
print(tab)
print('Ready !!')
如您所见,它打印出-- NEW TABLE --,但tab 变量是一个空数组。
tab 变量应该由每个表的学士学位、硕士和护理实践博士下的数据组成。
【问题讨论】:
标签: python html python-2.7 xpath web-scraping