【问题标题】:Beautiful soup check for tag in tag漂亮的汤检查标签中的标签
【发布时间】:2012-12-27 17:45:22
【问题描述】:

我正在使用 Beautiful Soup 4 来抓取页面。有一段我不想要的文字:

<p class="MsoNormal" style="text-align: center"><b>
                            <span lang="EN-US" style="font-family: Arial; color: blue">
                            <font size="4">1 </font></span>
                            <span lang="AR-SA" dir="RTL" style="font-family: Arial; color: blue">
                            <font size="4">&#1600;</font></span><span lang="EN-US" style="font-family: Arial; color: blue"><font size="4"> 
                            с&#1199;р&#1241; фати&#1211;&#1241;</font></span></b></p>

它的独特之处在于它有一个标签。我已经使用 findall() 来获取所有的

标签。所以现在我有一个 for 循环,例如:

for el in doc.findall('p'):
    if el.hasChildTag('b'):
        break;

可惜bs4没有“hasChildTag”功能

【问题讨论】:

    标签: python python-3.x screen-scraping beautifulsoup scraper


    【解决方案1】:

    应该也可以使用 css 选择器。

    http://www.crummy.com/software/BeautifulSoup/bs4/doc/#css-selectors

    soup.select("p b")
    

    【讨论】:

      【解决方案2】:
      for elem in soup.findAll('p'):
          if elem.findChildren('b'):
              continue #skip the elem with "b", and continue with the loop
          #do stuff with the elem
      

      【讨论】:

        猜你喜欢
        • 2015-05-03
        • 1970-01-01
        • 2016-06-02
        • 2019-09-06
        • 2020-03-15
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多