【问题标题】:beautifulsoup4 - grab Sibling element if Sibling presentbeautifulsoup4 - 如果 Sibling 存在,则获取 Sibling 元素
【发布时间】:2019-10-16 00:04:39
【问题描述】:

HTML 最常见的重复结构是:

  <p class="Standard">
   <span class="T3">
    it is possible for you
   </span>
  </p>

在这种情况下,我会抓住文字@9​​87654322@

偶尔(即并非总是),class="Standard" 的 &lt;p&gt; 有 class="P3" 的兄弟 &lt;p&gt;,如下所示:

  <p class="P3">
   (to ask a question in Spanish, you just use inflection)
  </p>

当这个&lt;p&gt; 的class="P3" 存在时,我想另外抓取其中的文本,例如这里我还要抢:(to ask a question in Spanish, you just use inflection)

我的问题是,鉴于这种结构:

<div>
...
  <p class="Standard">
   <span class="T3">
    it is possible for you
   </span>
  </p>

  <p class="Standard">
   <span class="T3">
    it is acceptable for me
   </span>
  </p>

  <p class="P3">
   (to ask a question in Spanish, you just use inflection)
  </p>
...
</div>

我怎样才能产生这样的输出:

it is possible for you
it is acceptable for me
(to ask a question in Spanish, you just use inflection)

目前,我已经做到了:

p_standards = soup.find_all("p", class_ = "Standard")

for p_standard in p_standards:
    p_english = p_standard.find("span", class_="T3")
    print(p_english.contents[0])

我得到的输出是:

it is possible for you
it is acceptable for me

【问题讨论】:

    标签: web-scraping beautifulsoup


    【解决方案1】:

    使用这个:

    Python 代码:

    from bs4 import BeautifulSoup
    import re
    text = '''
    <div>
    
      <p class="Standard">
       <span class="T3">
        it is possible for you
       </span>
      </p>
    
      <p class="Standard">
       <span class="T3">
        it is acceptable for me
       </span>
    
      </p>
      <p class="P3">
       (to ask a question in Spanish, you just use inflection)
      </p>
    
    </div>
    '''
    soup = BeautifulSoup(text,features='html.parser')
    p_standards = soup.find_all("p", class_ = "Standard")
    
    for p_standard in p_standards:
        p_english = p_standard.find('span',attrs={'class':'T3'})
        nextSibling = p_standard.find_next_sibling()
        print(p_english.text)
        if(nextSibling.attrs['class'][0] == 'P3' and nextSibling.name == 'p'):
          print(nextSibling.text)
    

    演示: Here

    说明:

    • 为了在 find_next_sibling's 中获取 class 值 返回的元素我必须搜索实例的变量 它的自我,因为没有文档在官方网站上提到它 所以我打印了nextSibling.__dict__.keys()
    • 0 索引是因为类属性的类型是数组

    【讨论】:

    • 在您的示例中,html 中有一个错字。 &lt;p class="P3"&gt; 元素不应嵌套在第二个 &lt;p class="Standard"&gt; 元素内。它是它的兄弟。
    • 如果您有时间,我将不胜感激,如果您能澄清您为什么需要这样做.attrs['class'][0]。请问[0]有什么意义?谢谢。 (此处的 BeautifulSoup4 文档是否包含此内容:crummy.com/software/BeautifulSoup/bs4/doc?)
    • 已添加说明部分,希望对您有所帮助
    【解决方案2】:

    我认为使用 css 或语法和相邻的兄弟组合器来执行此操作更有效

    from bs4 import BeautifulSoup as bs
    
    html = '''
    <div>
    ...
      <p class="Standard">
       <span class="T3">
        it is possible for you
       </span>
      </p>
    
      <p class="Standard">
       <span class="T3">
        it is acceptable for me
       </span>
      </p>
    
      <p class="P3">
       (to ask a question in Spanish, you just use inflection)
      </p>
    ...
    </div>
    '''
    soup = bs(html, 'lxml')
    items = [i.text.strip() for i in soup.select('.Standard, .Standard + .P3')]
    print(items)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2018-02-02
      • 1970-01-01
      • 1970-01-01
      • 2018-06-15
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多