【问题标题】:python: Replace/substitute all whole-word match in a stringpython:替换/替换字符串中的所有全字匹配
【发布时间】:2016-09-26 06:19:07
【问题描述】:

假设我的字符串是"#big and #small, #big-red, #big-red-car and #big"

如何使用re.sub(), re.match(), etc. 将一个标签替换为一个单词?

例如,所有#bigs 都必须更改为BIG,但#big-red 和#big-red-car 不应受到影响。

【问题讨论】:

  • 你不需要正则表达式。
  • 只需使用str的replace方法即可。
  • 更新了我的问题中的字符串。 @Keatinge 的建议不起作用,因为它将“#big-red”替换为“BIG-red”,这是不可取的。
  • @Keatinge,我尝试了空格技巧,但如果#big 位于字符串的末尾,那不会失败吗?或者后面是否有逗号或句号。

标签: python regex django string


【解决方案1】:

让我们定义你的字符串:

>>> s = "#big and #small, #big-red, #big-red-car and #big"

现在,让我们做你的替换:

>>> import re
>>> re.sub(r'#big([.,\s]|$)', r'#BIG\1', s)
'#BIG and #small, #big-red, #big-red-car and #BIG'

正则表达式#big([.,\s]|$) 将匹配所有#big 后跟句点、逗号、空格或或行尾的字符串。如果在#big 之后还有其他您认为可以接受的字符,您应该将它们添加到正则表达式中。

另类

如果我们想要更高级一点,可以使用前瞻断言(?=...) 来确保#big 后面的内容是可以接受的:

>>> re.sub(r'#big(?=[.,\s]|$)', r'#BIG', s)
'#BIG and #small, #big-red, #big-red-car and #BIG'

使用句点和逗号的测试

为了测试当#big 有"a comma or period after it" 时这是否正常工作,让我们创建一个新字符串:

>>> s = "#big and #big, #big. #small, #big-red, #big-red-car and #big"

然后,让我们测试一下:

>>> re.sub(r'#big(?=[.,\s]|$)', r'#BIG', s)
'#BIG and #BIG, #BIG. #small, #big-red, #big-red-car and #BIG'

【讨论】:

    【解决方案2】:

    此信息是一类单向边界技巧。

    使用否定向后/向前看断言,
    在特定方向内,它会让字符串的BEGIN/END匹配,
    但不允许其他人匹配。

    这导致了一些有趣的组合场景
    一个类中的否定结构,涵盖了无限的范围
    字符,但允许您排除其中的一些单独字符
    那个范围。

    使用的典型构造是否定类。

    \D - 非数字类
    \S - 非空白类
    \W - 非单词类
    \PP - 非标点属性类
    @987654325 @ - 非字母属性类

    因为它们被用在否定断言中,反之实际上是
    正在寻找的人物。

    \d, \s, \w, \pP, \pL分别

    力量来自于它们可以在内部组合的事实
    class 用于戏剧效果。

    如果单个字符被添加到一个类中,它们会被排除在外,不允许使用。
    实际上,它创建了类减法。

    创建类时的规则是:

    • 类你想要的字符,插入它是负数(即\D、\PP等)
    • 个别字符你不想要,照常插入(即\n、=等)
      这可以用作类减法。

    减法示例:(?![\S\r\n]) 将是一个前瞻边界,它需要
    只有水平空白,在某些引擎中,表示为
    \h 构造。


    在您的示例中,边界将是这样的。

    (?<![\S\PP-])#big(?![\S\PP-])

    打破它

     (?<!            # Boundary - Behind direction
          [\S\PP-]   # Need all whitespace and punctuation, but not the '-'
     )
     \#big
     (?!             # Boundary - Ahead direction
          [\S\PP-]   # Need all whitespace and punctuation, but not the '-'
     )
    

    添加到类中的每个文字字符实际上都排除了
    它来自匹配。

    这称为类减法。


    测试用例

    输入#big and #small, #big, #big, #big-red, #big-red-car and #big

    输出

     **  Grp 0 -  ( pos 0 , len 4 ) 
    #big  
    
     **  Grp 0 -  ( pos 17 , len 4 ) 
    #big  
    
     **  Grp 0 -  ( pos 23 , len 4 ) 
    #big  
    
     **  Grp 0 -  ( pos 56 , len 4 ) 
    #big  
    

    基本上,只匹配#big 和#small、#big、#big、#big-red、#big-red-car 和#big

    【讨论】:

      猜你喜欢
      • 2022-11-19
      • 1970-01-01
      • 1970-01-01
      • 2020-06-23
      • 2011-08-29
      • 2017-06-03
      • 2020-01-11
      • 1970-01-01
      • 2014-08-01
      相关资源
      最近更新 更多