【发布时间】:2016-02-22 13:35:39
【问题描述】:
我是新手,我试图用 python 编写一个蜘蛛,但我无法得到我需要的部分,我不知道哪里出了问题。
我从整个 html 文件中挑选出我需要的部分,如下所示。我尝试使用 RegEx </tr[/s]*?>(<tr[\s]*?>.*?</tr[\s]*?>)<tr[/s]*?>,但我什么也没得到。有没有人可以帮我解决这个问题?
PS。在使用 findall 收集信息之前,我已经使用 sub 删除了所有 \n 和 \r。
提前致谢。
</tr>
<tr >
<td style="border-bottom:3px solid #000000;" colspan='4' rowspan='1' >
<!-- START OBJECT-CELL -->
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='left' />
<col align='right' />
<tr>
<td align='left' bgcolor='#C0C0C0'><font color='#000000'>AE1PGA/L1/01</font></td>
<td align='right' bgcolor='#C0C0C0'><font color='#000000'>3-4</font></td>
</tr>
</table>
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='center' />
<tr>
<td align='center' bgcolor='#C0C0C0'><font color='#000000'>Programming And Algorithms</font></td>
</tr>
</table>
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='left' />
<tr>
<td align='left' bgcolor='#C0C0C0'><font color='#000000'></font></td>
</tr>
</table>
<!-- END OBJECT-CELL -->
</td>
<td style="border-bottom:3px solid #000000;" colspan='4' rowspan='1' >
<!-- START OBJECT-CELL -->
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='left' />
<col align='right' />
<tr>
<td align='left' bgcolor='#C0C0C0'><font color='#000000'>AE1MCS/L1/01</font></td>
<td align='right' bgcolor='#C0C0C0'><font color='#000000'>3-5, 7-15</font></td>
</tr>
</table>
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='center' />
<tr>
<td align='center' bgcolor='#C0C0C0'><font color='#000000'>Mathematics For Computer Scientists</font></td>
</tr>
</table>
<table bgcolor='#C0C0C0' cellspacing='0' border='0' width='100%'>
<col align='left' />
<tr>
<td align='left' bgcolor='#C0C0C0'><font color='#000000'>SEB-432+</font></td>
</tr>
</table>
<!-- END OBJECT-CELL -->
</td>
</tr>
<tr >
【问题讨论】:
-
Argh 不要自己解析 HTML,为此使用一个库:beautifulsoup4: pypi.python.org/pypi/beautifulsoup4(尤其是正则表达式对于这项工作来说是个糟糕的工具)
-
@RvdK 我试过bs4,但也不能正常工作,所以我必须自己做
-
bs4 不适合你怎么办?
-
你到底想从这个 HTML 中提取什么?您的预期输出是什么?
-
如果你使用
soup.find_all('tr')之类的东西可能会发生这种情况,但在 bs4 中过滤<tr>s 比使用正则表达式要容易得多
标签: python regex beautifulsoup