【问题标题】:Regular expression for class using Beautifulsoup使用 Beautifulsoup 的类的正则表达式
【发布时间】:2015-09-09 08:23:21
【问题描述】:

我正在使用 BeautifulSoup 来轻松抓取。

我发现网页中有超过 5 个 div 我想废弃。他们的名字不同,但有规律。

这些 div 是:

divnewthing
divnew
divnewstring

等

所以模式是divnew* 一种正则表达式。

我正在使用:

soup.find('div', {"class": "divnew"})

目前。

我想以某种方式使用正则表达式。有谁能帮帮我吗?

【问题讨论】:

    标签: python html regex beautifulsoup html-parsing


    【解决方案1】:

    是的,您也可以传递regular expression pattern:

    soup.find('div', {"class": re.compile("^divnew")})
    

    或者,一个函数,检查一个类名是否以divnew开头:

    soup.find('div', {"class": lambda x: x and x.startswith("divnew"))})
    

    或者,使用CSS selector:

    soup.select("div[class^=divnew]")
    

    【讨论】:

      猜你喜欢
      • 2017-06-30
      • 1970-01-01
      • 1970-01-01
      • 2016-05-01
      • 2015-10-11
      • 2014-12-10
      • 2012-09-01
      • 1970-01-01
      • 2021-07-20
      相关资源
      最近更新 更多