【问题标题】:how to avoid comments while extracting text from web pages in Python如何在 Python 中从网页中提取文本时避免评论
【发布时间】:2015-05-24 21:52:13
【问题描述】:

我正在尝试仅提取 网页中的文本,但我遇到了一些问题,例如没有写在页面中但它们是用 cmets 代码编写的文本,例如:“包括页脚”,“sidebar.php end”等。另外,我真的不想要的东西也来了。这是我用于测试用例的链接,即:

1) http://ai-depot.com/articles/the-easy-way-to-extract-useful-text-from-arbitrary-html

2) http://www.tutorialspoint.com/cplusplus/index.htm

3)http://www.cplusplus.com/doc/tutorial/program_structure/

(这样我就可以确保我的代码从任何页面中提取文本)

这是我遇到麻烦的代码:

import urllib
from bs4 import BeautifulSoup
url = "http://ai-depot.com/articles/the-easy-way-to-extract-useful-text-from-arbitrary-html/" 
html = urllib.urlopen(url).read()
soup = BeautifulSoup(html)

for script in soup(["script", "style","a","p","li","<!-->","small","<div id=\"footer\">","<div id=\"footer\">","<div id=\"bottom\">"]):
    script.extract()    

text = soup.findAll(text=True)
for p in text:
    print unicode(p)
fo = open('file.txt', 'w')
fo.seek(0, 2)
fo.writelines( unicode(p) )
fo.close()

在此代码中,我使用了 1 号链接,当我在该页面上执行 "inspect element" 时,我发现该代码中有很多 cmets,并且此代码也在提取它们。所以请帮忙.....

【问题讨论】:

  • 您的第一个链接是一篇文章,专门介绍了一种在抓取网站时减少 cmets 数量的方法。
  • 那我现在该怎么办?我是 pyhton 的新手。
  • 阅读...很多。这就是这里的每个知道任何事情的人都学到了他们所知道的东西。获得必要的技能和信息以实现您的目标可能需要数月或数年的时间,但这是一段宝贵的旅程,您将在此过程中学到很多其他技能。

标签: python


【解决方案1】:

当您的代码遇到正则表达式作为注释匹配的行时,一种方法是使用正则表达式去除/跳过 cmets。

或者,您也可以使用 HTML 解析器。 Python 在其标准库中内置了一个。

https://docs.python.org/2/library/htmlparser.html

【讨论】:

    猜你喜欢
    • 2015-09-28
    • 2020-09-03
    • 2018-04-10
    • 1970-01-01
    • 2021-09-15
    • 1970-01-01
    • 2014-07-24
    • 2017-02-24
    • 2019-03-11
    相关资源
    最近更新 更多