【问题标题】:Deleting rows according to the absence of a word in factor in R dataframe根据R数据框中缺少一个单词删除行
【发布时间】:2020-11-30 11:07:17
【问题描述】:

我有一个包含文本和作者的数据框。我只需要清理一个因子级别的数据,以保留存在一个单词的所有行。这是一个小例子:

author(factor)   text

John             Pear Plum

Mary             Pear Apple Banana Grapes

Mike             Grapes Apple Peach

John             Banana Pear Apple

John             Apple Melon 

这是我想要获得的结果,删除 John 没有提到 Apple 一词的每一行:

author(factor)   text

Mary             Pear Apple Banana Grapes

Mike             Grapes Apple Peach

John             Banana Pear Apple

John             Apple Melon 

这是我尝试过的:

df$author%in% "John"[!grepl("Apple", df$text, ignore.case = T),,drop = FALSE]

作为回应,我收到一条错误消息:

  Error in "John"[!grepl("Apple", df$text, ignore.case = T),  : 
  incorrect number of dimensions

我查看了有关子集数据的建议,但找不到与我的情况相似的任何内容。任何帮助表示赞赏。

【问题讨论】:

    标签: r dataframe data-cleaning grepl


    【解决方案1】:

    这行得通吗:

    library(dplyr)
    library(stringr)
    df %>% filter(!(author == 'John' & !str_detect(text, 'Apple')))
    # A tibble: 4 x 2
      author text                    
      <chr>  <chr>                   
    1 Mary   Pear Apple Banana Grapes
    2 Mike   Grapes Apple Peach      
    3 John   Banana Pear Apple       
    4 John   Apple Melon        
    

    使用的数据:

    df
    # A tibble: 5 x 2
      author text                    
      <chr>  <chr>                   
    1 John   Pear Plum               
    2 Mary   Pear Apple Banana Grapes
    3 Mike   Grapes Apple Peach      
    4 John   Banana Pear Apple       
    5 John   Apple Melon          
    

    【讨论】:

    • @Jess,奇怪,你的数据格式和你上面分享的一样吗?
    • 对不起,我在重写代码时犯了一个小错误!现在可以了,非常感谢!
    猜你喜欢
    • 2021-10-23
    • 2018-04-04
    • 2019-02-04
    • 2021-08-13
    • 2016-05-01
    • 2020-10-01
    • 2020-11-13
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多