【问题标题】:Working with broken HTML + BeautifulSoup处理损坏的 HTML + BeautifulSoup
【发布时间】:2018-01-20 04:28:10
【问题描述】:

我有一些非常糟糕的 HTML,长话短说,它阻止我使用普通的嵌套 <table>, <tr>, <td> 结构,这样可以很容易地重建表格。

这是一个带有行号的 sn-p 供参考:

1      <td valign="top">   <!-- closing </td> should be on 6 -->
2      <font face="arial" size="1">
3       <center>
4        06-30-95
5       </center>
6       <tr valign="top">
7        <td>
8         <center>
9          <font ,="" arial,="" face="arial" sans="" serif"="" size="1">
10          1382
11          <p>
12           (23)
13          </p>
14         </font>
15        </center>
16       </td>
17       <td>
18        <font ,="" arial,="" face="arial" sans="" serif"="" size="1">
19         <center>
20          06-18-14
21         </center>
22        </font>
23       </td>
24      </tr>
25    </td>    <!-- this should should be on 6 -->

trs 内的tds 内trs 内的嵌套没有任何方案,并且与未封闭的标签相结合以启动。 HTML 树与它的结构呈现方式完全不同。 (在这种情况下,我想技术上没有 missing 结束标记,但页面的实际呈现清楚地表明不应该嵌套 tds。)

但是,在这种情况下可以按照以下规则进行操作:

  • 对于任何&lt;td&gt;,在其关闭&lt;/td&gt;之前后跟一个开口&lt;td&gt;,(即任何嵌套的td)假设后一个开口&lt;td&gt;(第7行)作为第一个开口的闭包(第 1 行);
  • 否则,只需像往常一样抓住(打开、关闭)&lt;td&gt; ... &lt;/td&gt; 标签(其中开启者和关闭者之间没有&lt;td&gt;;例如上面的第 17 行和第 23 行。

这里想要的结果是这样的:

['06-30-95', '1382\n(23)', '06-18-14']

BeautifulSoup 如何解决这个问题?我会展示一个尝试,但是通过文档和一些源代码进行了挑选,但根本没有找到太多。

目前这将解析为:

html = """
<td valign="top">
 <font face="arial" size="1">
  <center>
   06-30-95
  </center>
  <tr valign="top">
   <td>
    <center>
     <font ,="" arial,="" face="arial" sans="" serif"="" size="1">
      1382
      <p>
       (23)
      </p>
     </font>
    </center>
   </td>
   <td>
    <font ,="" arial,="" face="arial" sans="" serif"="" size="1">
     <center>
      06-18-14
     </center>
    </font>
   </td>
  </tr>
</td>
"""

from bs4 import BeautifulSoup, SoupStrainer

strainer = SoupStrainer('td')
soup = BeautifulSoup(html, 'html.parser', parse_only=strainer)
[tag.text.replace('\n', '') for tag in soup.find_all('td')]

['   06-30-95        1382             (23)            06-18-14     ',
 '      1382             (23)      ',
 '      06-18-14     ']

我对这个结果的问题不是空格;这是子字符串的重复。似乎我需要从最里面的标签向上递归工作,弹出每个标签并向外工作。但我不得不猜测还有更多的内置功能可以处理缺少的结束标签(handle_endtag 从 BeautifulSoup 构造函数中脱颖而出?)。

【问题讨论】:

  • 解析成什么?

标签: python python-3.x beautifulsoup


【解决方案1】:

对于非常损坏的 HTML,有两种方法可以解决此问题。首先是在最里面的嵌套级别找到最一致的打开/关闭标签集,并且只使用第一个。在这个有限的例子中,看起来&lt;center&gt; 标签将满足这一点。考虑以下几点:

>>> from bs4 import BeautifulSoup
>>> soup = BeautifulSoup(html, 'html.parser')
>>> [t.find('center').text.strip() for t in soup.find_all('td')]
['06-30-95', '1382\n      \n       (23)', '06-18-14']

或者,使用lxml 代替(正如documentation 将其列为一种方法)实际上总体上可能效果更好:

>>> soup2 = BeautifulSoup(html, 'lxml')
>>> [t.text.strip() for t in soup2.find_all('td')]
['06-30-95', '1382\n      \n       (23)', '06-18-14']

此线程中还介绍了其他方法:Fast and effective way to parse broken HTML?

【讨论】:

    【解决方案2】:

    试试这个。它将获取您请求的输出:

    from bs4 import BeautifulSoup
    
    soup = BeautifulSoup(content, 'html5lib')
    item = [' '.join(items.text.split()) for items in soup.select("center")]
    print(item)
    

    输出:

    ['06-30-95', '1382 (23)', '06-18-14']
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2012-07-27
      • 2011-05-23
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-12-08
      相关资源
      最近更新 更多