【问题标题】:XPath keep track of data from every tableXPath 跟踪每个表中的数据
【发布时间】:2016-05-02 10:00:42
【问题描述】:

我一直在研究一个网站:

http://directory.ccnecommunity.org/reports/rptAccreditedPrograms_New.asp?sort=institution

我需要从各个大学提取硕士下的数据。

您可能注意到并非每所大学都有硕士数据,所以我需要跟踪它。

在这种情况下如何跟踪数据?

到目前为止,我的 python 和 XPATH 代码:

import __future__
from lxml import html
import requests
from bs4 import BeautifulSoup

page = requests.get('http://directory.ccnecommunity.org/reports/rptAccreditedPrograms_New.asp?sort=institution')

soup = str(BeautifulSoup(page.content, 'html.parser'))

tree = html.fromstring(soup)

for table in tree.xpath('//table[@width="95%" and @align="center" and @class="center"]'):
    print('-- NEW TABLE -- \n')
    tab = table.xpath('.//table[@width="260px"]/tr/td[@style="width: 100%;"]/text()')
    print(tab)

print('Ready !!')

如您所见,它打印出-- NEW TABLE --,但tab 变量是一个空数组。

tab 变量应该由每个表的学士学位、硕士和护理实践博士下的数据组成。

【问题讨论】:

    标签: python html python-2.7 xpath web-scraping


    【解决方案1】:

    试试:

    for table in tree.xpath('(//tr[ td[span="Baccalaureate"] or td[contains(span,"Master")] ]/ancestor::tr[1])'):
      print('-- NEW TABLE -- \n')
      tab = table.xpath('.//table[@width="260px"]/tr/td[@style="width: 100%;"]/text()')
      print(tab)
    

    【讨论】:

    • 它就像一个魅力。只是帮助我理解这一点。你为什么放:ancestor::tr[1]?以及为什么将td 作为tr 的索引,例如tr[ td ... ]
    • tr[ td ... ] 找到包含您要查找的数据的 tr。 ancestor::tr[1] 找到这个上面的下一个 tr。这个 tr 包含你在一所大学的所有数据。
    • 但是为什么它不像 tr/td[ ... ] 那样工作?为什么 td 必须是 tr 的索引?正如您所见,tab 的变量 xpath 不同...为什么?
    • 因为tr[ td ... ](带有和谓词的tr表达式)找到tr但tr/td[ ... ]找到td(tr/td现在是表达式)。而且第一个祖先不是你要找的那个。
    【解决方案2】:

    您可以使用以下 xpath 来提取 Master 的数据。

    //span[contains(text(),'Master')]/parent::td[1]
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-06-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-07-10
      • 1970-01-01
      • 2015-05-16
      • 1970-01-01
      相关资源
      最近更新 更多