【问题标题】:How to scrape elements that immediately follows a certain element?如何抓取紧跟在某个元素之后的元素?
【发布时间】:2015-12-27 08:17:08
【问题描述】:

我有一个如下所示的 Html 文档:

<div id="whatever">
  <a href="unwanted link"></a>
  <a href="unwanted link"></a>
  ...
  <code>blah blah</code>
  ...
  <a href="interesting link"></a>
  <a href="interesting link"></a>
  ...
</div>

我只想抓取紧跟code 标签的链接。如果我这样做 soup.findAll('a') 它会返回所有超链接。

如何让 BS4 在特定的 code 元素之后开始抓取?

【问题讨论】:

    标签: python beautifulsoup


    【解决方案1】:

    试试soup.find_all_next():

    >>> tag = soup.find('div', {'id': "whatever"})
    >>> tag.find('code').find_all_next('a')
    [<a href="interesting link"></a>, <a href="interesting link"></a>]
    >>> 
    

    它类似于soup.find_all(),但它会找到标签之后的所有标签。


    如果您想删除&lt;code&gt; 之前的&lt;a&gt; 标签,我们有一个名为soup.find_all_previous() 的函数:

    >>> tag.find('code').find_all_previous('a')
    [<a href="unwanted link"></a>, <a href="unwanted link"></a>]
    
    >>> for i in tag.find('code').find_all_previous('a'):
    ...     i.extract()
    ...     
    ... 
    <a href="unwanted link"></a>
    <a href="unwanted link"></a>
    
    >>> tag
    <div id="whatever">
    
    
      ...
      <code>blah blah</code>
      ...
      <a href="interesting link"></a>
    <a href="interesting link"></a>
      ...
    </div>
    >>> 
    

    那就是:

    1. 查找&lt;code&gt;标签之前的所有&lt;a&gt;标签。
    2. 使用 soup.extract() 和 for 循环删除它们。

    【讨论】:

    • 谢谢,我不知道这个功能。我想知道,有没有一种简单的方法来 extract() 或从 DOM 中删除出现在 code 元素之前的元素?
    • 太棒了!我需要消除其他元素,但您的代码提供了一个良好的开端!
    • @masroore:好吧,记住还有soup.findPrevious() 和soup.findNext(),如果你只需要找到一个元素而不是所有那个标签。查看the document了解更多详情。
    【解决方案2】:

    执行此操作的更简单的方法是将一串 css 选择器传递给 .select() 方法并使用 decompose 删除链接。在这里,您需要使用所谓的 General Sibling Selector ~ 来选择 所有 与 code 同级的锚点:code ~ a

    soup = BeautifulSoup('''<div id="whatever">
          <a href="unwanted link"></a>
          <a href="unwanted link"></a>
          ...
          <code>blah blah</code>
          ...
          <a href="interesting link"></a>
          <a href="interesting link"></a>
          ...
          </div>''', 
         'lxml'
    )
    
    for link in soup.select('code ~ a'):
        link.decompose()     
    
    print(soup)
    

    产生:

    <html><body><div id="whatever">
    <a href="unwanted link"></a>
    <a href="unwanted link"></a>
      ...
      <code>blah blah</code>
      ...
    
    
      ...
    </div></body></html>
    

    删除code 标记后所有链接的另一种方法是通过find_all 方法迭代列表返回以查找文档中的所有“代码”标记,并为每个标记使用find_all_next,它给出你是所有下一个a标签的列表。然后迭代列表并使用decompose 从树中删除标签,然后完全销毁它及其内容。

    演示

    In [85]: from bs4 import BeautifulSoup
    
    In [86]: soup = BeautifulSoup('''<div id="whatever">
       ....:   <a href="unwanted link"></a>
       ....:   <a href="unwanted link"></a>
       ....:   ...
       ....:   <code>blah blah</code>
       ....:   ...
       ....:   <a href="interesting link"></a>
       ....:   <a href="interesting link"></a>
       ....:   ...
       ....: </div>''', 'lxml')
    
    In [87]: for code in soup.find_all('code'):
       ....:     for link in code.find_all_next('a'):
       ....:         link.decompose()
       ....:         
    
    In [88]: soup
    Out[88]: 
    <html><body><div id="whatever">
    <a href="unwanted link"></a>
    <a href="unwanted link"></a>
      ...
      <code>blah blah</code>
      ...
    
    
      ...
    </div></body></html>
    

    【讨论】:

      猜你喜欢
      • 2020-05-10
      • 2021-07-02
      • 1970-01-01
      • 1970-01-01
      • 2023-03-27
      • 2014-09-26
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多