【发布时间】:2015-05-24 21:52:13
【问题描述】:
我正在尝试仅提取 网页中的文本,但我遇到了一些问题,例如没有写在页面中但它们是用 cmets 代码编写的文本,例如:“包括页脚”,“sidebar.php end”等。另外,我真的不想要的东西也来了。这是我用于测试用例的链接,即:
1) http://ai-depot.com/articles/the-easy-way-to-extract-useful-text-from-arbitrary-html
2) http://www.tutorialspoint.com/cplusplus/index.htm
3)http://www.cplusplus.com/doc/tutorial/program_structure/
(这样我就可以确保我的代码从任何页面中提取文本)
这是我遇到麻烦的代码:
import urllib
from bs4 import BeautifulSoup
url = "http://ai-depot.com/articles/the-easy-way-to-extract-useful-text-from-arbitrary-html/"
html = urllib.urlopen(url).read()
soup = BeautifulSoup(html)
for script in soup(["script", "style","a","p","li","<!-->","small","<div id=\"footer\">","<div id=\"footer\">","<div id=\"bottom\">"]):
script.extract()
text = soup.findAll(text=True)
for p in text:
print unicode(p)
fo = open('file.txt', 'w')
fo.seek(0, 2)
fo.writelines( unicode(p) )
fo.close()
在此代码中,我使用了 1 号链接,当我在该页面上执行 "inspect element" 时,我发现该代码中有很多 cmets,并且此代码也在提取它们。所以请帮忙.....
【问题讨论】:
-
您的第一个链接是一篇文章,专门介绍了一种在抓取网站时减少 cmets 数量的方法。
-
那我现在该怎么办?我是 pyhton 的新手。
-
阅读...很多。这就是这里的每个知道任何事情的人都学到了他们所知道的东西。获得必要的技能和信息以实现您的目标可能需要数月或数年的时间,但这是一段宝贵的旅程,您将在此过程中学到很多其他技能。
标签: python