【问题标题】:Extract the last word between | |提取 | 之间的最后一个单词|
【发布时间】:2021-07-10 12:23:23
【问题描述】:

我有以下数据集

> head(names$SAMPLE_ID)
[1] "Bacteria|Proteobacteria|Gammaproteobacteria|Pseudomonadales|Moraxellaceae|Acinetobacter|"
[2] "Bacteria|Firmicutes|Bacilli|Bacillales|Bacillaceae|Bacillus|"                            
[3] "Bacteria|Proteobacteria|Gammaproteobacteria|Pasteurellales|Pasteurellaceae|Haemophilus|" 
[4] "Bacteria|Firmicutes|Bacilli|Lactobacillales|Streptococcaceae|Streptococcus|"             
[5] "Bacteria|Firmicutes|Bacilli|Lactobacillales|Streptococcaceae|Streptococcus|"             
[6] "Bacteria|Firmicutes|Bacilli|Lactobacillales|Streptococcaceae|Streptococcus|" 

我想将|| 之间的最后一个单词提取为一个新变量,即

Acinetobacter
Bacillus
Haemophilus

我尝试过使用

library(stringr)
names$sample2 <-   str_match(names$SAMPLE_ID, "|.*?|")

【问题讨论】:

  • 简单路线:vapply(strsplit(names$SAMPLE_ID, "|", fixed = TRUE), tail, "", 1)
  • 或者你们中的一些人不喜欢打字(或效率)然后sapply(strsplit(x, "\\|"), tail, 1)

标签: regex r stringr


【解决方案1】:

我们可以使用

library(stringi)
stri_extract_last_regex(v1, '\\w+')
#[1] "Acinetobacter"

数据

v1 <- "Bacteria|Proteobacteria|Gammaproteobacteria|Pseudomonadales|Moraxellaceae|Acinetobacter|"

【讨论】:

  • 似乎stringi包在某些方面优于stringr。
【解决方案2】:

仅使用基础 R:

myvar <- gsub("^..*\\|(\\w+)\\|$", "\\1", names$SAMPLE_ID)

【讨论】:

    【解决方案3】:
    ^.*\\|\\K.*?(?=\\|)
    

    使用\K从最后一场比赛中删除其余部分。参见演示。也使用perl=T

    https://regex101.com/r/fM9lY3/45

    x <- c("Bacteria|Firmicutes|Bacilli|Lactobacillales|Streptococcaceae|Streptococcus|",
           "Bacteria|Firmicutes|Bacilli|Lactobacillales|Streptococcaceae|Streptococcus|" )
    
    unlist(regmatches(x, gregexpr('^.*\\|\\K.*?(?=\\|)', x, perl = TRUE)))
    # [1] "Streptococcus" "Streptococcus"
    

    【讨论】:

      【解决方案4】:

      结局就是你想要的[^|]+(?=\|$)

      根据@RichardScriven:

      Which in R would be regmatches(x, regexpr("[^|]+(?=\\|$)", x, perl = TRUE)

      【讨论】:

        【解决方案5】:

        在这种情况下,您也可以使用包“stringr”。代码如下:

        v<- "Bacteria| Proteobacteria|Gammaproteobacteria|Pseudomonadales|Moraxellaceae|Acinetobacter|"

        v1&lt;- str_replace_all(v, "\\|", " ")

        word(v1,-2)

        这里我使用 v 作为字符串。基本原理是将|全部替换为空格,然后通过函数word()得到字符串中的最后一个单词。

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2018-12-31
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多