【问题标题】:How to get value between two different tags using beautiful soup?如何使用漂亮的汤在两个不同的标签之间获得价值?
【发布时间】:2017-07-22 02:40:05
【问题描述】:

我需要在下面的代码 sn-p 中提取结束标记和
标记之间存在的数据:

<td><b>First Type :</b>W<br><b>Second Type :</b>65<br><b>Third Type :</b>3</td>

我需要的是:W, 65, 3

但问题是这些值也可以是空的,比如-

<td><b>First Type :</b><br><b>Second Type :</b><br><b>Third Type :</b></td>

如果存在这些值,我想获取这些值,否则为空字符串

我尝试使用 nextSibling 和 find_next('br') 但它返回了

 <br><b>Second Type :</b><br><b>Third Type :</b></br></br>

和

<br><b>Third Type :</b></br>

如果标签之间不存在值(W、65、3)

</b> and <br> 

我需要的是,如果这些标签之间没有任何内容,它应该返回一个空字符串。

【问题讨论】:

  • 嘿,next_sibling 最终对我来说很好:)

标签: python beautifulsoup html-parsing


【解决方案1】:

我会使用 &lt;b&gt; 标签 by &lt;/b&gt; 标签策略,查看他们的 next_sibling 包含什么类型的信息。

我会检查他们的next_sibling.string 是否不是None,并相应地附加列表:)

>>> html = """<td><b>First Type :</b><br><b>Second Type :</b>65<br><b>Third Type :</b>3</td>"""

>>> soup = BeautifulSoup(html, "html.parser")
>>> b = soup.find_all("b")
>>> data = []
>>> for tag in b:
        if tag.next_sibling.string == None:
            data.append(" ")
        else:
            data.append(tag.next_sibling.string)
>>> data 
[' ', u'65', u'3'] # Having removed the first string

希望这会有所帮助!

【讨论】:

  • 这正是我想要的。谢谢!!
【解决方案2】:

我会搜索td 对象,然后使用regex 模式过滤您需要的数据,而不是在find_all 方法中使用re.compile。

像这样:

import re
from bs4 import BeautifulSoup

example = """<td><b>First Type :</b>W<br><b>Second Type :</b>65<br><b>Third 
Type :</b>3</td>
<td><b>First Type :</b><br><b>Second Type :</b>69<br><b>Third Type :</b>6</td>"""

soup = BeautifulSoup(example, "html.parser")

for o in soup.find_all('td'):
    match = re.findall(r'</b>\s*(.*?)\s*(<br|</br)', str(o))
    print ("%s,%s,%s" % (match[0][0],match[1][0],match[2][0]))

此模式查找&lt;/b&gt; 标记和&lt;br&gt; 或&lt;/br&gt; 标记之间的所有文本。 &lt;/br&gt; 标签是在将汤对象转换为字符串时添加的。

这个例子输出:

W,65,3

,69,6

只是一个例子,如果其中一个正则表达式匹配为空,您可以更改为返回一个空字符串。

【讨论】:

  • 它只是我拥有的大型 html 文件的一个小 sn-p。使用 re 作为数据容器的结构是否有效,即 和
    标签之间的信息在整个 html 文件中保持相同?
  • 是的,它很有效,如果用自己的代码替换重新匹配,你不会错过任何东西。
【解决方案3】:
In [5]: [child for child in soup.td.children if isinstance(child, str)]
Out[5]: ['W', '65', '3']

这些文本和标签是 td 的孩子,您可以使用 contents(list) 或 children(generator) 访问它们

In [4]: soup.td.contents
Out[4]: 
[<b>First Type :</b>,
 'W',
 <br/>,
 <b>Second Type :</b>,
 '65',
 <br/>,
 <b>Third Type :</b>,
 '3']

那么你可以通过测试是否是str的实例来获取文本

【讨论】:

    【解决方案4】:

    我认为这可行:

    from bs4 import BeautifulSoup
    html = '''<td><b>First Type :</b>W<br><b>Second Type :</b>65<br><b>Third Type :</b>3</td>'''
    soup = BeautifulSoup(html, 'lxml')
    td = soup.find('td')
    string = str(td)
    list_tags = string.split('</b>')
    list_needed = []
    for i in range(1, len(list_tags)):
        if list_tags[i][0] == '<':
            list_needed.append('')
        else:
            list_needed.append(list_tags[i][0])
    print(list_needed)
    #['W', '65', '3']
    

    因为你想要的值总是在标签的末尾,这样很容易捕捉到它们,不需要重新。

    【讨论】:

      猜你喜欢
      • 2016-06-02
      • 1970-01-01
      • 2020-08-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-03-30
      • 1970-01-01
      相关资源
      最近更新 更多