【问题标题】:Remove a combination of characters in a string python删除字符串python中的字符组合
【发布时间】:2021-03-04 04:35:16
【问题描述】:

正则表达式新手警报请温柔。我有这样的字符串:

sent = 'The type of vehicle FRFR7800 is the fastest'

我想删除重复出现的子字符串“FR”。所以字符串应该是:

sent = 'The type of vehicle FR7800 is the fastest'.

我想我已经花了两个多小时阅读/尝试re 的教程和groupby 的更pythonic 方式,我真的想不通。我还搜索了类似的问题,大多数结果都涵盖了我重复相同字符的情况,例如有像 'dddddaaaaaggggg' 之类的字符串。其中一些有帮助,但我最终删除了所有出现的 'FR'。

例如我试过:

sent = re.sub(r'FC{1}', '', sent)
sent = re.sub(r'FC|', '', sent)

这些完全消除了“FR”的出现。当我将其更改为:

sent = re.sub(r'FC{2}', '', sent)

什么也没发生,字符串仍然重复出现“FR”。

有人可以帮我或给我一个提示吗?

【问题讨论】:

  • @VishalSingh 您已将 FR 硬编码为子字符串,我怀疑这是 OP 想要的。
  • 已修复@TimBiegeleisen regex101.com/r/Jg99O9/1
  • @Vishal Singh 谢谢你的回答。我还发现我不知道如何在 python 中语法。我尝试了它并删除了所有出现的情况。我试过了:send = re.sub(r'(\w{2})\1', '', sent)

标签: python python-3.x regex string


【解决方案1】:
import re

sent = "The type of vehicle FRFR7800 is the fastest"
regex = r"(\w{2})\1"

print(re.sub(regex, r"\g<1>", sent))

【讨论】:

  • 谢谢!这解决了它。没有意识到在替换字符串参数( r"\g" )中我也可以使用正则表达式。
  • 这很好,但要小心你真正理解你将要处理的文本,因为它替换了所有像这样的对,而且这些对存在于英语中。 IE。使用"The papal cocoa-puff dispenser of hippopotamus-shaped vehicle FRFR7800 is in crisis"查看结果
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2019-07-26
  • 1970-01-01
  • 2013-06-11
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-08-29
相关资源
最近更新 更多