【问题标题】:beautifulsoup 4 multiple class values regexbeautifulsoup 4 多个类值正则表达式
【发布时间】:2016-01-23 10:49:12
【问题描述】:

我得到了以下标签:

<span class="one two-l three-project">not interested in</span>
<span class="one two-xl">interesting 1</span>
<span class="one two-l">interesting 2</span>

如何选择 2 个较低的跨度而不选择第一个?

我试过了:

html.find_all('span', attrs={'class' : re.compile(r'(one two-)'}) // doesn't select anything

html.find_all('span', attrs={'class' : re.compile(r'(?!three-project)')}) // does select all

html.find_all('span', attrs={'class' : 'one two-xl')}) // doesn't select the 3rd one

html.select('span.one.two-xl') // doesn't select the 3rd one

任何想法都非常感谢:-)

【问题讨论】:

  • 不幸的是,bs4 将“one two-l three-project”解释为不是一个完整的大字符串,而是 3 个字符串(one、two-l 和 three-project)。因此“one two-x?l(?!three-project)”根本不匹配。
  • Beautifulsoup 不是很擅长这种东西,你想要什么lxml.de,文中有什么独特的地方可以用吗?

标签: python regex class beautifulsoup


【解决方案1】:

BeautifulSoup 功能的限制意味着选择标签 with 一些类,但 没有 其他类。我不会使用正则表达式,而是使用列表推导来遍历所有跨度并使用if 语句选择相关的跨度。

您可以执行以下操作:

from bs4 import BeautifulSoup

content = '''
<span class="one two-l three-project">not interested in</span>
<span class="one two-xl">interesting 1</span>
<span class="one two-l">interesting 2</span>
'''

soup = BeautifulSoup(content)
desired_tags = [tag for tag in soup.find_all('span') if 'three-project' not in tag.attrs['class']]
print(desired_tags)

输出

[<span class="one two-xl">interesting 1</span>,
 <span class="one two-l">interesting 2</span>]

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2015-09-09
    • 1970-01-01
    • 2021-07-20
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多