【发布时间】:2013-08-10 00:38:25
【问题描述】:
学习 Python,我正在尝试制作一个没有任何 3rd 方库的网络爬虫,这样我的过程就不会简化,而且我知道自己在做什么。我浏览了几个在线资源,但所有这些都让我对某些事情感到困惑。
html 看起来像这样,
<html>
<head>...</head>
<body>
*lots of other <div> tags*
<div class = "want" style="font-family:verdana;font-size:12px;letter-spacing:normal"">
<form class ="subform">...</form>
<div class = "subdiv1" >...</div>
<div class = "subdiv2" >...</div>
*lots of other <div> tags*
</body>
</html>
我希望爬虫提取 <div class = "want"...>*content*</div> 并将其保存到 html 文件中。
我对如何解决这个问题有一个非常基本的想法。
import urllib
from urllib import request
#import re
#from html.parser import HTMLParser
response = urllib.request.urlopen("http://website.com")
html = response.read()
#Some how extract that wanted data
f = open('page.html', 'w')
f.write(data)
f.close()
【问题讨论】:
-
我可以理解我不想使用一个为你做所有事情的网络抓取库......但你可能想考虑使用
BeautifulSoup只是用于解析部分。如果 HTML 都是现代且有效的,那么您应该可以在 stdlib 中使用它,但如果您需要处理古怪的现实页面,BS 会让您的生活更轻松。 (即使对于简单的情况,它也稍微简单一些,但这没什么大不了的。)
标签: python web-scraping extract extraction