【问题标题】:Subset rows in an R dataframe based on partial match of multiple strings基于多个字符串的部分匹配的 R 数据框中的子集行
【发布时间】:2020-02-25 22:46:59
【问题描述】:

我认为没有人问过这个确切的问题 - 很多关于基于一个值(即x[grepl("some string", x[["column1"]]),])的子集的东西,但不是多个值/字符串。

这是我的数据示例:

#create sample data frame
data = data.frame(id = c(1,2,3,4), phrase = c("dog, frog, cat, moose", "horse, bunny, mouse", "armadillo, cat, bird,", "monkey, chimp, cow"))

#convert the `phrase` column to character string (the dataset I'm working on requires this)
data$phrase = data$phrase

#list of strings to remove rows by
remove_if = c("dog", "cat")

这将给出一个如下所示的数据集:

  id                phrase
1  1 dog, frog, cat, moose
2  2   horse, bunny, mouse
3  3 armadillo, cat, bird,
4  4    monkey, chimp, cow

我想删除第 1 行和第 3 行(因为第 1 行包含“狗”,第 3 行包含“猫”),但保留第 2 行和第 4 行。

  id                phrase
1  2   horse, bunny, mouse
2  4    monkey, chimp, cow

换句话说,我想对data 进行子集化,使其只有(标题和)第 2 行和第 4 行(因为它们既不包含“狗”也不包含“猫”)。

谢谢!

【问题讨论】:

    标签: r grepl


    【解决方案1】:

    我们可以在paste将“remove_if”转换为单个字符串之后使用grepl 和subset

    subset(data, !grepl(paste(remove_if, collapse="|"), phrase))
    #    id              phrase
    #2  2 horse, bunny, mouse
    #4  4  monkey, chimp, cow
    

    【讨论】:

    • 在真实数据集上按预期工作 - 谢谢!
    【解决方案2】:

    使用grep

    > data[grep(paste0(remove_if, collapse = "|"), data$phrase, invert = TRUE), ]
      id              phrase
    2  2 horse, bunny, mouse
    4  4  monkey, chimp, cow
    

    【讨论】:

      【解决方案3】:

      如果您想将其与dplyr 和stringr 混合使用:

      library(stringr)
      library(dplyr)
      
      data %>%
        filter(str_detect(phrase, paste(remove_if, collapse = "|"), negate = TRUE))
      #   id              phrase
      # 1  2 horse, bunny, mouse
      # 2  4  monkey, chimp, cow
      

      【讨论】:

        【解决方案4】:
        data[!grepl(paste0("(^|, )(", paste0(remove_if, collapse = "|"), ")(,|$)"), data$phrase),]
        
        # id                    phrase
        #  2 caterpillar, bunny, mouse
        #  4        monkey, chimp, cow
        

        在这个例子中构造的正则表达式是"(^|, )(dog|cat)(,|$)",以避免匹配包含“cat”或“dog”但实际上不是确切的单词的单词,例如'毛毛虫'

        【讨论】:

          【解决方案5】:

          另一种方式(也许不是最好的方式):

          data[-unique(unlist(sapply(c(remove_if),function(x){grep(x,data$phrase)}))),]
            id              phrase
          2  2 horse, bunny, mouse
          4  4  monkey, chimp, cow
          

          【讨论】:

            猜你喜欢
            • 2019-09-02
            • 2017-01-21
            • 1970-01-01
            • 2016-10-18
            • 2016-04-15
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 2020-11-25
            相关资源
            最近更新 更多