【发布时间】:2019-11-03 07:18:37
【问题描述】:
我有一个 BeautifulSoup 段落作为字符串。我想使用正则表达式替换字符串中出现的 p (开始)和 /p (结束)标签,因为有像
这样的实例 <p class="section-para">We would be happy to hear from you, Please
fill in the form below or mail us your requirements on<br/><span
class="text-red">contact@xyz.com</span></p>
但我不能使用泛型
^< *>$
因为我想要 strong、b 和 h1,h1..h6 标记用于不同的目的。
我只知道 RegEx 的基础知识,但不知道如何制作和使用。 有人可以帮我制作“包含”、“排除”(如果有的话)。我怎样才能为这个问题做一个,我怎样才能用简单的 ''
代替def formatting(string):
this=['<h1>','</h1>','<h2>','</h2>','<h3>','</h3>','<h4>','</h4>','<h5>','</h5>','<h6>','</h6>','<b>','</b>','<strong>','</strong>']
with_this=['\nh1 Tag:','\n','\nh2 Tag:','\n''\nh3 Tag:','\n''\nh4 Tag:','\n''\nh5 Tag:','\n''\nh6 Tag:','\n','\Bold:','\n''\nBold:','\n']
for i in range(len(this)):
if this[i] in string:
string=string.replace(this[i],with_this[i])
return(string)
我已经为 h1,2...6 标签使用了字符串的替换功能。任何帮助将不胜感激。
【问题讨论】:
-
如果你已经努力使用漂亮的汤来解析你的 html,那么不要使用正则表达式。从不推荐使用正则表达式来解析/替换 html
-
嗨!通常,使用 RegEx 解析 HTML 被认为是一件傻事。见stackoverflow.com/a/1732454/3543867。您可能会遇到很多边缘情况。处理 html 的一个常见参考是 BeautifulSoup 模块/库 - stackoverflow.com/questions/11709079/parsing-html-using-python
-
虽然据说 Chuck Norris 能够在 html 上使用正则表达式 ...
-
虽然你不能编写一个只匹配有效 html 的正则表达式,但无论如何尝试这样做都会将可怕的 Zalgo 召唤到我们的现实中,只是试图匹配一个
标签实际上可能是可能的,因为您不能将
嵌套在另一个
中(或者至少,您不应该这样做)。
标签: python html regex python-3.x beautifulsoup