【发布时间】:2017-12-14 02:18:19
【问题描述】:
我正在尝试删除 (html 标签) 中的文本并将结果写入新文件。例如,一行文本可能是:
< asdf> Text <here>more text< /asdf >
因此程序将写入输出文件:“Text more text”,不包括那些在 html 标签内的内容。
这是我目前的尝试:
import urllib.request
data=urllib.request.urlopen("some website").read()
text1=data.decode("utf-8")
import re
def asd(text1):
x=re.compile("<>")
y=re.sub(x,"",text1)
file1=open("textfileoutput.txt","w")
file1.write(y)
return y
asd(text1)
好像没有写干净的版本,还是有标签的。感谢您的帮助。
【问题讨论】:
-
您的正则表达式只会匹配“”。我建议像BeautifulSoup Grab Visible Webpage Text 这样的解决方案。
-
你是对的,用这个替换一行来修复它: x=re.compile(r"]+>") 程序现在可以工作了。谢谢。
-
如果标签在某处包含 > 怎么办?正如 alecxe 指出的,尝试用正则表达式解析 HTML 通常不是最好的。
标签: python html regex html-parsing