【问题标题】:Issue with negative lookbehind in RR中的负面回顾问题
【发布时间】:2020-07-07 09:17:42
【问题描述】:

我有这组句子:

w <- c("so i said er well it would n't surprise me if it could bloody talk",  # quote marker
        "we got fifteen, well thirteen minutes",                              
        "well she brought a pie and she brought some er punch round",         
        "so your dad said well have n't i been soft ?",                       # quote marker
        "And he went [pause] well I can't feel any. ",                        # quote marker
        "I goes well they'll improve the grant to start off with",            # quote marker
        "so with the chips as well this is about one sixty .",                
        "well we 're not all the same are we , but") 

所有字符串都包含单词well。我对well 充当引号标记的那些字符串感兴趣,如said、goes 和went 的出现所示。使用 positive lookbehind 我可以匹配这些句子:

grep("(?<=said|goes|went).*well", w, value = T, perl = T)
[1] "so i said er well it would n't surprise me if it could bloody talk"
[2] "so your dad said well have n't i been soft ?"                      
[3] "And he went [pause] well I can't feel any. "                       
[4] "I goes well they'll improve the grant to start off with"

我遇到的问题是 negative 向后查找以匹配那些 'well' 是 not 引号标记的字符串 not 起作用。例如,这匹配所有内容:

grep("(?<!said|goes|went).*well", w, value = T, perl = T)
[1] "so i said er well it would n't surprise me if it could bloody talk" # not match
[2] "we got fifteen, well thirteen minutes"                              # match
[3] "well she brought a pie and she brought some er punch round"         # match    
[4] "so your dad said well have n't i been soft ?"                       # not match         
[5] "And he went [pause] well I can't feel any. "                        # not match             
[6] "I goes well they'll improve the grant to start off with"            # not match         
[7] "so with the chips as well this is about one sixty ."                # match      
[8] "well we 're not all the same are we , but"                          # match

为什么不正确匹配?如何更改才能正确匹配?

提前致谢!

【问题讨论】:

  • 您不只是想在积极的后视中反转您的匹配吗? grep(..., invert = TRUE)
  • 我对这个选项很熟悉,但在帖子中我特别有兴趣深入挖掘负面的后视。不过还是谢谢。

标签: r regex regex-lookarounds


【解决方案1】:

发生这种情况是因为(?&lt;!said|goes|went) 匹配字符串中的位置,该位置不是立即 以在lookbehind 中定义的字符串之前。 .* 然后尽可能多地匹配除换行符以外的任何 0+ 字符,然后匹配 well。有很多这样的有效职位。

最简单的方法是匹配said、goes 或went 出现在well 之前的那些字符串并跳过它们,然后在所有其他上下文中匹配well:

\b(?:said|goes|went)\b.*\bwell\b(*SKIP)(*F)|\bwell\b

请参阅regex demo。

警告:如果您使用^(?!.*\b(?:said|goes|went)\b).*\bwell\b 之类的解决方案,当said、goes 或went 出现之后 @987654336,您可能会得到误报@。

模式详情

  • \b(?:said|goes|went)\b.*\bwell\b(*SKIP)(*F) - 一个完整的单词:said,goes 或 went,然后是尽可能多的任何 0 个或多个字符,然后是一个完整的单词well,在找到此匹配项后,将其删除并正则表达式引擎开始在当前失败的位置寻找匹配项
  • | - 或
  • \bwell\b - 一个完整的词well。

查看R demo:

grep("\\b(?:said|goes|went)\\b.*\\bwell\\b(*SKIP)(*F)|\\bwell\\b", w, value = TRUE, perl = TRUE)
# [1] "we got fifteen, well thirteen minutes"                     
# [2] "well she brought a pie and she brought some er punch round"
# [3] "so with the chips as well this is about one sixty ."       
# [4] "well we 're not all the same are we , but"    

【讨论】:

    猜你喜欢
    • 2013-08-03
    • 1970-01-01
    • 1970-01-01
    • 2016-05-19
    • 2016-09-07
    • 1970-01-01
    • 1970-01-01
    • 2022-11-10
    • 2018-06-29
    相关资源
    最近更新 更多