【问题标题】:How to subset a dataframe where a column 'contains' the contents of a vector如何对列“包含”向量内容的数据框进行子集化
【发布时间】:2014-06-24 00:27:42
【问题描述】:

我有一个这样的数据框 DF:

ID  name     description
1   Jen      brooklyn new york cqre center
2   Chris    santa monica barbers
3   Jeff     chicago blah blah
4   Steve    groomers are here
5   Mary     chicken time
6   John     and here we go 

我想删除包含城市名称的行。我有一个包含城市名称(cityVec)的向量,但这不起作用。

f

【问题讨论】:

  • 试试DF[-grep(paste(cityVec,collapse="|"), DF$description, ignore.case=TRUE),]
  • 我收到以下错误: grep 错误(paste(cityVec, collapse = "|"), DF$description, ignore.case = TRUE) : 正则表达式在此语言环境中无效
  • 你的数据中可能有这样的字符"é",请确认
  • 我认为问题在于它是一个向量。 Grep 适用于单个字符串,但包含字符串的向量似乎不起作用
  • 在下面查看我的答案,它与@jdharrison 的答案相同,grep 确实适用于字符串向量,您能否发布str(DF) 的输出,grep 将不起作用如果DF$description 是因素

标签: r


【解决方案1】:

从 jdharrison 答案中分叉数据输入,grep 的工作方式与您在字符串向量上看到的一样

description = list(c("brooklyn", "new york", "cqre center")
                   , c("santa monica", "barbers")
                   , c("chicago", "blah blah")
                   , c("groomers", "are here")
                   , c("chicken time")
                   , c("and", "here we go"))

DF <- data.frame(ID = 1:6, name = letters[1:6], description = I(description))

cityVec <- c("chicago", "new york", "paris", "santa monica") 
myDF[-grep(paste(cityVec,collapse="|"), myDF$description, ignore.case=TRUE),]
#  ID name  description
#4  4    d groomers....
#5  5    e chicken time
#6  6    f and, her....

【讨论】:

  • | 的美妙用法。
  • 这两种方法都会导致内存错误。它可能适用于小型数据集,但大型数据框似乎对这种方法存在问题。
  • 您在使用这些方法时遇到的内存错误是什么,您能否更新您的帖子。有大内存处理经验的人可能会帮助你
  • 我相信它毕竟是格式错误的向量。花了一段时间才在这么大的向量中找到坏字符。谢谢大家。
【解决方案2】:

我假设您的 data.frame 有 description 的列表,如果没有,请使用 dput。您可以使用 sapply 搜索每个元素并检查它是否包含城市

description = list(c("brooklyn", "new york", "cqre center")
                   , c("santa monica", "barbers")
                   , c("chicago", "blah blah")
                   , c("groomers", "are here")
                   , c("chicken time")
                   , c("and", "here we go"))

myDF <- data.frame(ID = 1:6, name = letters[1:6], description = I(description))

cityVec <- c("chicago", "new york", "paris", "santa monica")                
> myDF[sapply(myDF$description, function(x){!any(x%in%cityVec)}), ]
ID name  description
4  4    d groomers....
5  5    e chicken time
6  6    f and, her....

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2012-10-27
    • 2013-04-22
    • 1970-01-01
    • 1970-01-01
    • 2020-11-05
    • 1970-01-01
    • 2012-12-30
    • 1970-01-01
    相关资源
    最近更新 更多