【问题标题】:Select a html a tag with specified display content选择一个带有指定显示内容的html标签
【发布时间】:2017-09-19 13:46:21
【问题描述】:

我是 scrapy 的新手,已经为这个问题苦苦挣扎了好几个小时。
我需要抓取一个页面,它的来源看起来像这样:

 <tr class="odd">
          <td class="pfama_PF02816">Pfam</td>
          <td><a href="http://pfam.xfam.org/family/Alpha_kinase">Alpha_kinase</a></td>
          <td>1389</td>
          <td>1590</td>
          <td class="sh" style="display: none">21.30</td>
        </tr>  

我需要获取tr.odd标签的信息,当且仅当a标签具有“Alpha_kinase”值
我可以获取所有这些内容(包括“Alpha_kinase”、1389、1590 和许多其他值),然后处理输出以仅获取“Alpha_kinase”,但这种方法将非常脆弱和丑陋。目前我必须这样做:
positions = response.css('tr.odd td:not([class^="sh"]) td a::text').extract() 然后做一个for循环来检查。
是否有任何条件(如上面的td.not)表达式可以放入 response.css 来解决我的问题?

提前致谢。任何建议将不胜感激!

【问题讨论】:

  • 你有什么限制吗?你试过正则表达式吗?
  • @rjustin 我没有。我认为有点像 td: 不像我上面指定的那样,但很难找到这样的表达方式。你对我有什么建议吗?

标签: python html scrapy


【解决方案1】:

您可以使用另一个选择器:response.xpath 从 html 中选择元素,

并使用 xpath contains 函数过滤文本。

>>> response.xpath("//tr[@class='odd']/td/a[contains(text(),'Alpha_kinase')]")
[<Selector xpath="//tr[@class='odd']/td/a[contains(text(),'Alpha_kinase')]" data='<a href="http://pfam.xfam.org/family/Alp'>]

【讨论】:

  • 不错的建议。我要试试这个
【解决方案2】:

我假设页面上有多个这样的tr 元素。如果是这样,我可能会这样做:

# get only rows containing 'Alpha_kinase' in link text
for row in response.xpath('//tr[@class="odd" and contains(./td/a/text(), "Alpha_kinase")]'):
    # extract all the information
    item['link'] = row.xpath('./td[2]/a/@href').extract_first()
    ...
    yield item

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-05-31
    • 1970-01-01
    • 2011-09-28
    • 1970-01-01
    • 2017-10-31
    相关资源
    最近更新 更多