【问题标题】:Regular expression to extract all text after "--!!" in R dplyr正则表达式提取“--!!”之后的所有文本在 R dplyr
【发布时间】:2020-04-30 17:38:18
【问题描述】:

我正在尝试在 R 中使用 dplyr 在以下示例中由变量 name 的某些实例过滤的数据帧中的变量字符串之后提取子字符串。我试图将所需的结果传递给一个名为income_rent 的新变量。

我是正则表达式的新手。我的尝试是:

income_cashrent <- v18 %>% 
filter(str_detect(name, "B25122")) %>% 
mutate(income_rent = str_extract(label, "[^--!!]*$"))

但是,我得到了结果: Error in stri_extract_first_regex(string, pattern, opts_regex = opts(pattern)) : Syntax error in regexp pattern. (U_REGEX_RULE_SYNTAX)

name的前四行是:

Estimate!!Total
Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000
Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent
Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent!!Less than $100

期望的结果是:

[not sure how to indicate an empty result here]
Less than $10,000
Less than $10,000!!With cash rent
Less than $10,000!!With cash rent!!Less than $100

到目前为止,我一直无法调试这个,参考堆栈上的其他正则表达式示例。任何指导都将受到欢迎。提前谢谢大家!

【问题讨论】:

    标签: r regex dplyr


    【解决方案1】:
    regmatches(vec, gregexpr("(?<=--!!).*", vec, perl = TRUE))
    # [[1]]
    # character(0)
    # [[2]]
    # [1] "Less than $10,000"
    # [[3]]
    # [1] "Less than $10,000!!With cash rent"
    # [[4]]
    # [1] "Less than $10,000!!With cash rent!!Less than $100"
    

    如果您从这里unlist,您会注意到您“丢失”了第一个条目,不确定这是否有问题。

    unlist(regmatches(vec, gregexpr("(?<=--!!).*", vec, perl = TRUE)))
    # [1] "Less than $10,000"                                
    # [2] "Less than $10,000!!With cash rent"                
    # [3] "Less than $10,000!!With cash rent!!Less than $100"
    

    如果这是个问题,那么

    vecout <- regmatches(vec, gregexpr("(?<=--!!).*", vec, perl = TRUE))
    unlist(replace(vecout, lengths(vecout) < 1, NA))
    # [1] NA                                                 
    # [2] "Less than $10,000"                                
    # [3] "Less than $10,000!!With cash rent"                
    # [4] "Less than $10,000!!With cash rent!!Less than $100"
    

    (或者您也可以替换为""。)


    dplyr 管道中:

    tibble(vec = c("Estimate!!Total",
    # "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000",
    # "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent",
    # "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent!!Less than $100")) %>%
      mutate(out = regmatches(vec, gregexpr("(?<=--!!).*", vec, perl = TRUE)), out = replace(out, lengths(vecout) < 1, NA), out = unlist(out))
    + + # A tibble: 4 x 2
    #   vec                                             out                           
    #   <chr>                                           <chr>                         
    # 1 Estimate!!Total                                 <NA>                          
    # 2 Estimate!!Total!!Household income in the past ~ Less than $10,000             
    # 3 Estimate!!Total!!Household income in the past ~ Less than $10,000!!With cash ~
    # 4 Estimate!!Total!!Household income in the past ~ Less than $10,000!!With cash ~
    

    数据:

    vec <- c("Estimate!!Total",
    "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000",
    "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent",
    "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent!!Less than $100")
    

    【讨论】:

    • 这本身就可以很好地工作,但我还没有设法让它与之前的过滤条件一起工作。不确定它是否不能嵌套在 dplyr 管道中。尝试了regmatches(v18[which(str_detect(v18$name, "B25122"))], gregexpr("(?&lt;=--!!).*", v18, perl = TRUE)),但这引发了错误。
    • 查看我的编辑,这正是列的代码,因为它是独立向量的代码,不确定你做了什么不同。
    【解决方案2】:

    我们可以使用str_extract to extract the characters after the pattern--!!` 使用正则表达式查找

    library(stringr)
    library(dplyr)
     v18 %>%        
         mutate(income_rent = str_extract(label, "(?<=--!!).*"))                                                                                                                                                label
    #1                                                                                                                                    Estimate!!Total
    #2                                 Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000
    #3                 Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent
    #4 Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent!!Less than $100
     #                                       income_rent
    #1                                              <NA>
    #2                                 Less than $10,000
    #3                 Less than $10,000!!With cash rent
    #4 Less than $10,000!!With cash rent!!Less than $100
    

    或者另一个选项是str_match

    v18$income_rent <-  str_match(v18$label, ".*--!!(.*)")[,2]
    

    数据

    v18 <- structure(list(label = c("Estimate!!Total", "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000", 
    "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent", 
    "Estimate!!Total!!Household income in the past 12 months (in 2018 inflation-adjusted dollars) --!!Less than $10,000!!With cash rent!!Less than $100"
    )), class = "data.frame", row.names = c(NA, -4L))
    

    【讨论】:

    • 感谢您的回答。然而,我应该在我的数据样本中更加具体:并非源变量中的所有条目在“--!!”之后都有那个“小于”字符串。我需要得到“--!!”之后的任何内容,所以不幸的是,这个解决方案对我不起作用。
    • @AbeBarranca 在这种情况下,只需将其更改为正则表达式环视即可。更新。谢谢
    • 现在完美运行。谢谢,阿克伦!正如我所说,我是正则表达式的新手,所以不知道环视是如何工作的。这非常有用。
    • @AbeBarranca 谢谢,我还更新了str_match 以便它可以直接提取元素而不是使用正则表达式环视
    • str_match 解决方案很感兴趣。您能否告知解决方案结束时[,2] 索引的作用?
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-06-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多