【问题标题】:Replacing the values in a string for something like <p *anything inside it*> to just empty '' or Nothing将字符串中的值替换为 <p *anything inside it*> 为空 '' 或 Nothing
【发布时间】:2019-11-03 07:18:37
【问题描述】:

我有一个 BeautifulSoup 段落作为字符串。我想使用正则表达式替换字符串中出现的 p (开始)和 /p (结束)标签,因为有像

这样的实例
    <p class="section-para">We would be happy to hear from you, Please 
    fill in the form below or mail us your requirements on<br/><span 
    class="text-red">contact@xyz.com</span></p> 

但我不能使用泛型

    ^< *>$

因为我想要 strong、b 和 h1,h1..h6 标记用于不同的目的。

我只知道 RegEx 的基础知识,但不知道如何制作和使用。 有人可以帮我制作“包含”、“排除”(如果有的话)。我怎样才能为这个问题做一个,我怎样才能用简单的 ''

代替
def formatting(string):
    this=['<h1>','</h1>','<h2>','</h2>','<h3>','</h3>','<h4>','</h4>','<h5>','</h5>','<h6>','</h6>','<b>','</b>','<strong>','</strong>']
    with_this=['\nh1 Tag:','\n','\nh2 Tag:','\n''\nh3 Tag:','\n''\nh4 Tag:','\n''\nh5 Tag:','\n''\nh6 Tag:','\n','\Bold:','\n''\nBold:','\n']

    for i in range(len(this)):
        if this[i] in string:
            string=string.replace(this[i],with_this[i])
    return(string)

我已经为 h1,2...6 标签使用了字符串的替换功能。任何帮助将不胜感激。

【问题讨论】:

标签: python html regex python-3.x beautifulsoup


【解决方案1】:

不清楚你到底想替换什么,但也许下面的内容会有所帮助,如果你需要的话,它可以让你用文本替换标签。相信您将能够进一步调整以使其达到您想要的效果。此外,您没有指定您正在使用的 BS 版本。我正在使用BS4。该函数将接受一个美丽的汤对象,一个要查找的标签,一个前缀 I.E 你想用什么替换开始标签和一个后缀,即你想用什么替换结束标签。

from bs4 import BeautifulSoup

def format_soup_tag(soup, tag, prefix, suffix):
    target_tag = soup.find(tag)
    target_tag.insert_before(prefix)
    target_tag.insert_after(suffix)
    target_tag.unwrap()

html = '<p class ="section-para">We would be happy to hear from you, <strong>Please fill in the form below</strong> or mail us your requirements on <br/><span class ="text-red" >contact@xyz.com</span></p>'
soup = BeautifulSoup(html, features="lxml")
print("###before modification###\n", soup, "\n")

format_soup_tag(soup, 'p', '\np tag: ', '\n')
print("###after p tag###\n", soup, "\n")

format_soup_tag(soup, 'strong', '\Bold: ', ' \Bold')
print("###after strong tag###\n", soup, "\n")

输出

###before modification###
 <html><body><p class="section-para">We would be happy to hear from you, <strong>Please fill in the form below</strong> or mail us your requirements on <br/><span class="text-red">contact@xyz.com</span></p></body></html> 

###after p tag###
 <html><body>
p tag: We would be happy to hear from you, <strong>Please fill in the form below</strong> or mail us your requirements on <br/><span class="text-red">contact@xyz.com</span>
</body></html> 

###after strong tag###
 <html><body>
p tag: We would be happy to hear from you, \Bold: Please fill in the form below \Bold or mail us your requirements on <br/><span class="text-red">contact@xyz.com</span>
</body></html> 

【讨论】:

  • stackoverflow.com/questions/56688597/… 。这是我正在处理的原始问题。我正在使用 BS4
  • 感谢您的帮助。我还想添加 h1,h2,h3 值,就像你为 strong 所做的那样,如果存在于段落中。
  • 该函数的编写方式很灵活,因此您可以根据需要添加更多行,例如format_soup_tag(soup, 'h1', '\nh1 tag: ', '\n'),您甚至可以使用前缀数组制作标签名称字典和后缀并遍历字典
【解决方案2】:

我希望我理解正确,如果我错了,请纠正我。你有类似的东西:

<p class="section-para">We would be happy to hear from you, Please 
fill in the form below or mail us your requirements on<br/><span 
class="text-red">contact@xyz.com</span></p>

并且想要类似的东西:

<p>We would be happy to hear from you, Please 
fill in the form below or mail us your requirements on<br/><span 
class="text-red">contact@xyz.com</span></p>

你可以这样做:

saved_content = re.search(
    '<p (.*?)>(?P<content>.*)</p>',
    your_string
).groupdict()

result = re.sub(
    r'<p (.*?)>(.*)</p>',
    f'<p>{saved_content.get("content")}</p>',
    your_string
)

请注意,我使用了仅在 Python 3.6 或更高版本中可用的 f 字符串。我希望它对您有所帮助,如果我误解了任何内容或有任何问题,请告诉我。祝你有美好的一天!

【讨论】:

  • 谢谢你的回答,有点像这样。但是我想做的原作是这样的-stackoverflow.com/questions/56688597/…
  • 如果这个答案是这个特定问题的正确答案,您应该将其标记为正确答案。我也会看看其他问题。
  • yuor 其他问题没有意义。为什么 p 标签内会有任何标题 (h1,h2...h6) 标签。就html树结构而言,标题标签基本上会关闭p标签。如果您查看常见问题解答,获得帮助的最佳方式是创建一个最低限度的完整示例来说明您的尝试。这意味着提供您输入的样本。显示您期望的输出并显示您尝试过的任何代码。
猜你喜欢
  • 2014-11-12
  • 2014-08-10
  • 1970-01-01
  • 1970-01-01
  • 2013-11-04
  • 2022-09-24
  • 1970-01-01
  • 2018-12-04
  • 2019-06-07
相关资源
最近更新 更多