【问题标题】:Filtering for multiple strings within the same column in r过滤r中同一列中的多个字符串
【发布时间】:2019-10-23 12:17:49
【问题描述】:

我的大型数据集(Groceries)中有一个包含字符数据(Fruits)的列,所有这些数据都是小写的,并且所有这些都没有标点符号。

看起来有点像这样:

# Groceries Data Frame
Id    Groceries$Fruits
1     apple orange banana lemon grapefruit
2     grapes tomato passion fruit
3     strawberry orange kiwi
4     lemon orange passion fruit grapefruit lime
5     lemon orange passion fruit grapefruit lime peach
  ...

我正在尝试从 Fruits 列中选择包含 5 种特定水果(橙子、酸橙、柠檬、葡萄柚和百香果)的所有行(其中有 3,320 行)。最初,我只对包含所有 5 种水果的行感兴趣,而没有其他水果。因此,这 5 行中唯一应该过滤/子集的行将是第 4 行。水果不必按任何特定顺序排列。

数据实际上是一个测试的答案,所以最终我有兴趣确定谁得到了 0/5 水果,谁得到了 1/5、2/5 等等......

到目前为止,我已经尝试了 2 种方法,但都无济于事。 首先我尝试使用 grep(),但结果数据框中没有存储任何行。

# 1st attempt with grep()
Correct fruits <- Groceries[grep("orange, lemon, lime, passion fruit, 
grapefruit", Groceries$Fruits), ]

然后我尝试使用 filter(),但所选行不只包含我正在寻找的 5 种水果,它会选择包含 5 种水果中的任何一种的所有行。

# 2nd attempt with filter
library(dplyr)
library(stringr)
CorrectFruits <- c("lemon", "orange", "passion fruit", "grapefruit", 
"lime")

filter <- Groceries %>%
  select(Id, Fruits) %>%
  filter(str_detect(tolower(Fruits), pattern = CorrectFruits))

我最初追求的结果是一个新的 DF,其中包含 Groceries 表中的所有列,但只有那些正确选择了所有 5 个水果的人的行。

接下来,选择相反的会很酷;所有 5 项都没有答对的人。

最后,我希望能够对那些获得正确特定比例的人进行子集化。 IE。第 1 行答对了 3 个,第 2 行只答对了 1 个,第 3 行只答对了 1 个。

任何帮助将不胜感激!

以下是一些列的示例:

# Groceries
Id   Age      Nationality    Colour question   Fruits question
1    26-35    Canadian       Red               apple orange banana lemon grapefruit
2    26-35    US             Blue              grapes tomato passion fruit
3    46-55    Canadian       Red               strawberry orange kiwi
4    55+      US             Red               lemon orange passion fruit grapefruit lime
5    36-45    British        Green             lemon orange passion fruit grapefruit lime peach

【问题讨论】:

  • 您能否提供一个可重现的 Grocery 数据集示例?部分数据会有所帮助。
  • 您好 Jim O,感谢您这么快回来,尽管数据集有数千行长,但我在底部添加了一些数据的简短示例。如果我可以添加任何有用的东西,请告诉我一些细节
  • 当然,我会试试的。但是为了更好地为您提供帮助,您的查询是否始终限于 5 个水果?
  • 是的,问题是“尽可能多地命名这 5 种水果”,上面列出了 5 种水果的 5 张单独图片
  • 如果您在 tidyverse 中工作,我建议不要将您的函数命名为 filter 以避免潜在的冲突。

标签: r string filter subset


【解决方案1】:

可能需要进一步澄清您打算如何处理所有 5 种水果和一些额外的答案,但这应该对您有所帮助。我用“passionfruit”替换了所有的“passionfruit”实例以使其更容易:

df$Fruits <- gsub("passion fruit", "passionfruit", df$Fruits)
CorrectFruits <- c("lemon", "orange", "passionfruit", "grapefruit", 
                   "lime")
df$Count <- str_count(df$Fruits, paste(CorrectFruits, collapse = '|'))
df$Count <- ifelse((df$Count == 5 & str_count(df$Fruits, '\\w+') > 5), 0, df$Count)

给了

ID                                          Fruits Count
1            apple orange banana lemon grapefruit     3
2                      grapes tomato passionfruit     1
3                          strawberry orange kiwi     1
4       lemon orange passionfruit grapefruit lime     5
5 lemon orange passionfruit grapefruit lime peach     0

第一行进行百香果替换,然后 str_count 计算 df$Fruit 中所有正确水果的出现次数。最后,如果所有 5 个水果都正确但有多余的,Count 重置为 0。

【讨论】:

  • 这真的很有帮助,非常感谢。如果您愿意,我有 2 次跟进... 1) 就额外的水果答案而言,您如何更改代码,而不是将他们的分数重置为 0,而是每增加一个水果就失去一分? 2)如果我想创建一个新的 df,保留 Groceries 数据框中的所有数据,但只保留得分为 5 的行,我该怎么做?
  • 我不明白第一个后续问题。请看下面我的回答。
【解决方案2】:

这是我看到别人的天才解决方案后的答案。

ID <- c(1:5)
Age <- c("26-35", "26-35", "46-55", "55+", "56-45")
Nationality <- c("Canadian", "US", "Canadian", "US", "British")
Color <- c("Correct", "Incorrect", "Incorrect", "Correct", "Correect")
Fruits <- c("pineapple", 
            "apple", 
            "apple orange kiwi fifth",
            "orange apple pineapple kiwi fifth",
            "pineapple orange apple fifth kiwi"
            )
df <- data.frame(ID, Age, Nationality, Color, Fruits)
df

heds1 的回复看起来很棒。但是,您要小心使用字符串精确,例如 grepl,因为它可能返回复合词。例如,考虑单词pineapple;它包含 pine 和 apple。注意这里搜索苹果返回菠萝。

filter(df, grepl("apple", Fruits))

  ID   Age Nationality     Color                            Fruits
1  1 26-35    Canadian   Correct                         pineapple
2  2 26-35          US Incorrect                             apple
3  3 46-55    Canadian Incorrect           apple orange kiwi fifth
4  4   55+          US   Correct orange apple pineapple kiwi fifth
5  5 56-45     British  Correect pineapple orange apple fifth kiwi

sumshyftw 提供的答案很棒。我喜欢我从 sumshyftw 那里学到一些东西。但是为了证明我的观点,无限制的字符串搜索可能会弄乱你的计数:

CorrectFruits <- c("apple")
df$Count <- str_count(df$Fruits, paste(CorrectFruits, collapse = '|'))
df$Count <- ifelse((df$Count == 5 & str_count(df$Fruits, '\\w+') > 5), 0, df$Count)
df

  ID   Age Nationality     Color                            Fruits Count
1  1 26-35    Canadian   Correct                         pineapple     1
2  2 26-35          US Incorrect                             apple     1
3  3 46-55    Canadian Incorrect           apple orange kiwi fifth     1
4  4   55+          US   Correct orange apple pineapple kiwi fifth     2
5  5 56-45     British  Correect pineapple orange apple fifth kiwi     2

请注意,尽管唯一正确的水果是苹果,但它仍将菠萝视为正确答案。为了克服这个问题,你想用\\b 包装你的话。

CorrectFruits <- c("\\bapple\\b")
df$Count <- str_count(df$Fruits, paste(CorrectFruits, collapse = '|'))
df$Count <- ifelse((df$Count == 5 & str_count(df$Fruits, '\\w+') > 5), 0, df$Count)
df

  ID   Age Nationality     Color                            Fruits Count
1  1 26-35    Canadian   Correct                         pineapple     0
2  2 26-35          US Incorrect                             apple     1
3  3 46-55    Canadian Incorrect           apple orange kiwi fifth     1
4  4   55+          US   Correct orange apple pineapple kiwi fifth     1
5  5 56-45     British  Correect pineapple orange apple fifth kiwi     1

R 不再将菠萝视为苹果。

但为了记录,sumshyftw 在我的示例中解决困难部分值得称赞:

CorrectFruits <- c("\\bapple\\b", "\\borange\\b", "\\bpineapple\\b", "\\bfifth\\b", "\\bkiwi\\b")
df$Count <- str_count(df$Fruits, paste(CorrectFruits, collapse = '|'))
df$Count <- ifelse((df$Count == 5 & str_count(df$Fruits, '\\w+') > 5), 0, df$Count)
df

  ID   Age Nationality     Color                            Fruits Count
1  1 26-35    Canadian   Correct                         pineapple     1
2  2 26-35          US Incorrect                             apple     1
3  3 46-55    Canadian Incorrect           apple orange kiwi fifth     4
4  4   55+          US   Correct orange apple pineapple kiwi fifth     5
5  5 56-45     British  Correect pineapple orange apple fifth kiwi     5

只显示所有五种水果的人:

df2 <- filter(df, df$Count == 5)
df2

  ID   Age Nationality    Color                            Fruits Count
1  4   55+          US  Correct orange apple pineapple kiwi fifth     5
2  5 56-45     British Correect pineapple orange apple fifth kiwi     5

【讨论】:

  • 没有考虑菠萝的问题!喜欢你的回答。
【解决方案3】:

这是使用带有目标关键字列表的grepl 的一种方法。

df <- structure(list(v1 = structure(1:4, .Label = c("row1", "row2", 
"row3", "row4"), class = "factor"), v2 = structure(c(2L, 4L, 
1L, 3L), .Label = c("another invalid row", "apple banana mandarin orange pear", 
"banana apple mandarin pear orange", "not a valid row"), class = "factor")), class = "data.frame", row.names = c(NA, 
-4L))

targets <- c("banana", "apple", "orange", "pear", "mandarin")
bool_df <- as.data.frame(sapply(targets, grepl, df$v2))
match_rows <- which(rowSums(bool_df) == 5)
df <- df[match_rows,]

然后您可以通过将5 更改为4 来更改match_rows 向量中的条件,例如四个水果匹配的4,等等。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2018-03-26
    • 2018-07-08
    • 2018-05-13
    • 2015-12-30
    • 1970-01-01
    • 2021-12-09
    • 1970-01-01
    相关资源
    最近更新 更多