【问题标题】:Python removing website html tags not workingPython删除网站html标签不起作用
【发布时间】:2017-12-14 02:18:19
【问题描述】:

我正在尝试删除 (html 标签) 中的文本并将结果写入新文件。例如,一行文本可能是:

< asdf> Text <here>more text< /asdf >

因此程序将写入输出文件:“Text more text”,不包括那些在 html 标签内的内容。

这是我目前的尝试:

import urllib.request

data=urllib.request.urlopen("some website").read()

text1=data.decode("utf-8")

import re

def asd(text1):

    x=re.compile("<>")

    y=re.sub(x,"",text1)

    file1=open("textfileoutput.txt","w")

    file1.write(y)

    return y

asd(text1)

好像没有写干净的版本,还是有标签的。感谢您的帮助。

【问题讨论】:

  • 您的正则表达式只会匹配“”。我建议像BeautifulSoup Grab Visible Webpage Text 这样的解决方案。
  • 你是对的,用这个替换一行来修复它: x=re.compile(r"]+>") 程序现在可以工作了。谢谢。
  • 如果标签在某处包含 > 怎么办?正如 alecxe 指出的,尝试用正则表达式解析 HTML 通常不是最好的。

标签: python html regex html-parsing


【解决方案1】:
x=re.compile("<>")

我不确定你为什么认为这个表达式会匹配 &lt; asdf&gt; 或 &lt; /asdf &gt;。

无论如何,使用正则表达式can rarely be justified 处理HTML。 为该任务使用更合适的工具 - HTML 解析器。

使用BeautifulSoup 的示例是unwrap() method:

In [1]: from bs4 import BeautifulSoup

In [2]: html = "<asdf>Text more text</asdf>"

In [3]: soup = BeautifulSoup(html, "html.parser")

In [4]: soup.asdf.unwrap()
Out[4]: <asdf></asdf>

In [5]: print(soup)
Text more text

【讨论】:

  • 对于一些关心性能的人来说,BeautifulSoup 真的很慢,甚至使用lxml 作为解析器。如果您的 html 文本格式正确并且您信任您的正则表达式,那么使用它就没有问题。
【解决方案2】:

只需将re.compile("&lt;&gt;") 替换为re.compile(r"&lt;[^&lt;&gt;]*&gt;") 就足够了

【讨论】:

  • 如果标签在某处包含 > 怎么办?
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2020-03-08
  • 2017-02-22
  • 2018-11-09
  • 1970-01-01
  • 1970-01-01
  • 2017-07-01
  • 1970-01-01
相关资源
最近更新 更多