【问题标题】:Matching partial ids in BeautifulSoup匹配 BeautifulSoup 中的部分 id
【发布时间】:2011-02-19 07:21:51
【问题描述】:

我正在使用 BeautifulSoup。我必须找到任何对 <div> 标签的引用,其 id 如下:post-#。

例如:

<div id="post-45">...</div>
<div id="post-334">...</div>

我试过了:

html = '<div id="post-45">...</div> <div id="post-334">...</div>'
soupHandler = BeautifulSoup(html)
print soupHandler.findAll('div', id='post-*')

如何过滤?

【问题讨论】:

  • 你用的是什么版本的 BeautifulSoup?

标签: python beautifulsoup


【解决方案1】:

您可以将函数传递给findAll:

>>> print soupHandler.findAll('div', id=lambda x: x and x.startswith('post-'))
[<div id="post-45">...</div>, <div id="post-334">...</div>]

或正则表达式:

>>> print soupHandler.findAll('div', id=re.compile('^post-'))
[<div id="post-45">...</div>, <div id="post-334">...</div>]

【讨论】:

  • AttributeError: 'NoneType' 对象没有属性 'startswith'
  • 我已经修复了AttributeError。
【解决方案2】:
soupHandler.findAll('div', id=re.compile("^post-$"))

我觉得很合适。

【讨论】:

  • 你为什么放$?我认为这不会像 OP 所希望的那样起作用。
【解决方案3】:

由于他要求匹配“post-#somenumber#”,因此最好精确

import re
[...]
soupHandler.findAll('div', id=re.compile("^post-\d+"))

【讨论】:

    【解决方案4】:

    这对我有用:

    from bs4 import BeautifulSoup
    import re
    
    html = '<div id="post-45">...</div> <div id="post-334">...</div>'
    soupHandler = BeautifulSoup(html)
    
    for match in soupHandler.find_all('div', id=re.compile("post-")):
        print match.get('id')
    
    >>> 
    post-45
    post-334
    

    【讨论】:

      猜你喜欢
      • 2017-07-31
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-07-01
      • 2015-07-29
      • 2013-11-03
      • 2010-12-14
      • 2016-01-08
      相关资源
      最近更新 更多