【问题标题】:Extracting an element whose id starts with a certain string using BeautifulSoup in python [duplicate]在python中使用BeautifulSoup提取id以某个字符串开头的元素[重复]
【发布时间】:2023-03-14 11:20:02
【问题描述】:

我正在尝试使用 BS4 进行一些网页抓取。

到目前为止,我已经提取了 <a> 使用

urls = [item for item in soup.select('h4 a')]

但是,我只想知道 ID 从哪个条目开始的 url。

<a href="http://www.sampleurl.com/static/welcome" id="entry_1">Lamborghini </a>

我尝试过item.id,但它不起作用。

我错过了什么?

【问题讨论】:

  • item.get('id')?
  • 是的,如果条件是“ID 开始 with 'entry'”urls = [item for item in soup.select('h4 a') if item.get("id", "")[:6] == "entry_"]

标签: python beautifulsoup


【解决方案1】:

将re 模块与id 一起使用。
方法如下:

from bs4 import BeautifulSoup
import re

if __name__ == "__main__":
    html = '<a href="http://www.sampleurl.com/static/welcome" id="entry_1">Lamborghini </a>'
    soup = BeautifulSoup(html, 'html.parser')

    print(soup.find('a', id=re.compile('^entry_')))

输出:

<a href="http://www.sampleurl.com/static/welcome" id="entry_1">Lamborghini </a>

【讨论】:

猜你喜欢
  • 2016-12-17
  • 1970-01-01
  • 1970-01-01
  • 2012-01-17
  • 1970-01-01
  • 2014-12-02
  • 2015-11-10
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多