【问题标题】:case_when() issue with evaluating multiple conditionscase_when() 评估多个条件的问题
【发布时间】:2021-07-23 23:46:52
【问题描述】:

我正在尝试检测字符串中是否存在特定的关键字和短语,如果它们存在,我想在新列中发布特定的数字。我的问题是某些字符串有多个关键字,但 case_when 只返回第一个匹配项。有没有办法解决这个问题,或者我应该使用 case_when 的替代方法?

ID<-c(1,2,3,4,5)
fruits<-c("banana apple orange", "apple orange", "orange", "orange apple", "nothing")
df<-data_frame(ID,fruits)
#I need to assign a random number to each fruit type

df %>% 
  mutate("Fruit Type"=case_when(
    grepl("banana",fruits)~34,
    grepl("apple",fruits)~45,
    grepl("orange",fruits)~88,
))

ID       fruits                  Fruit Type
1      banana apple orange           34
2      apple orange                  45
3      orange                        88
4      orange apple                  45
5      nothing                       NA

我希望它会像这样出来。

ID        fruits       fruit_type         
1   banana apple orange    34       
1   banana apple orange    45       
1   banana apple orange    88       
2   apple orange           45       
2   apple orange           88       
3   orange                 88       
4   orange apple           88       
4   orange apple           45       
5   nothing                NA

此外,有没有办法将其转换为长格式以使其看起来更像这样?

ID        fruits       fruit_type  fruit_type2  fruit_type3     
1   banana apple orange    34        45              88                                     
2   apple orange           45        88              NA                                     
3   orange                 88        NA              NA                             
4   orange apple           88        45              NA                                                     
5   nothing                NA        NA              NA         

【问题讨论】:

    标签: r dplyr


    【解决方案1】:

    这是separate、str_detect 和across 的另一种解决方案:

    library(dplyr)
    library(tidyr)
    library(stringr)
    
    df %>% 
      separate(fruits, c("fruit_type1", "fruit_type2", "fruit_type3"), remove = FALSE) %>% 
      mutate(across(contains("fruit_type"), ~case_when(
        str_detect(., "banana") ~ 34,
        str_detect(., "apple") ~ 45,
        str_detect(., "orange") ~ 88)
      ))
    

    输出:

         ID fruits              fruit_type1 fruit_type2 fruit_type3
      <dbl> <chr>                     <dbl>       <dbl>       <dbl>
    1     1 banana apple orange          34          45          88
    2     2 apple orange                 45          88          NA
    3     3 orange                       88          NA          NA
    4     4 orange apple                 88          45          NA
    5     5 nothing                      NA          NA          NA
    

    【讨论】:

      【解决方案2】:

      我们通过使用str_replace_all 使用命名向量更改字符串值来获得第一个输出,然后在空格处拆分“fruit_type”列以扩展数据(separate_rows),并更改“fruit_type”的类型到numeric类

      library(dplyr)
      library(tidyr)
      library(stringr)
      out <- df %>% 
          mutate(fruit_type = str_replace_all(fruits, 
           setNames(as.character(c(34, 45, 88)), c("banana", "apple", "orange")))) %>% 
          separate_rows(fruit_type) %>%
          mutate(fruit_type = as.numeric(fruit_type))
      

      -输出

      out
      # A tibble: 9 x 3
           ID fruits              fruit_type
        <dbl> <chr>                    <dbl>
      1     1 banana apple orange         34
      2     1 banana apple orange         45
      3     1 banana apple orange         88
      4     2 apple orange                45
      5     2 apple orange                88
      6     3 orange                      88
      7     4 orange apple                88
      8     4 orange apple                45
      9     5 nothing                     NA
      

      有了这个输出,我们可以用pivot_wider重塑为“宽”格式

      library(data.table)
      out %>% 
          mutate(rn = str_c('fruit_type', rowid(ID))) %>% 
          pivot_wider(names_from = rn, values_from = fruit_type)
      

      -输出

      # A tibble: 5 x 5
           ID fruits              fruit_type1 fruit_type2 fruit_type3
        <dbl> <chr>                     <dbl>       <dbl>       <dbl>
      1     1 banana apple orange          34          45          88
      2     2 apple orange                 45          88          NA
      3     3 orange                       88          NA          NA
      4     4 orange apple                 88          45          NA
      5     5 nothing                      NA          NA          NA
      

      【讨论】:

        【解决方案3】:

        你可以先把不同的水果放在不同的行里,然后用case_when-

        library(dplyr)
        library(tidyr)
        
        res <- df %>%
          separate_rows(fruits, sep = '\\s+') %>%
          mutate(Fruit_Type =case_when(
            grepl("banana",fruits)~34,
            grepl("apple",fruits)~45,
            grepl("orange",fruits)~88,
          ))
          
        res
        
        #     ID fruits  Fruit_Type
        #  <dbl> <chr>        <dbl>
        #1     1 banana          34
        #2     1 apple           45
        #3     1 orange          88
        #4     2 apple           45
        #5     2 orange          88
        #6     3 orange          88
        #7     4 orange          88
        #8     4 apple           45
        #9     5 nothing         NA
        

        要获得宽格式的数据,您可以这样做 -

        res %>%
          group_by(ID) %>%
          mutate(row = paste0('Fruit', row_number()), 
                 fruits = paste0(fruits, collapse = ' ')) %>%
          ungroup %>%
          pivot_wider(names_from = row, values_from = Fruit_Type)
        
        #    ID fruits              Fruit1 Fruit2 Fruit3
        #  <dbl> <chr>                <dbl>  <dbl>  <dbl>
        #1     1 banana apple orange     34     45     88
        #2     2 apple orange            45     88     NA
        #3     3 orange                  88     NA     NA
        #4     4 orange apple            88     45     NA
        #5     5 nothing                 NA     NA     NA
        

        【讨论】:

        • 这与我尝试做的非常接近,但在我的数据集中,除了水果名称之外,还有更多的单词,所以我无法按每个单词分开。例如,我的数据集中的每一行都说“有一个香蕉、一个苹果和一个橙子”。使用seperate_rows,当我只想要相关水果单词的一行时,这会为每个单词创建一个新行。有没有办法使用seperate_rows 只针对特定的单词?
        【解决方案4】:

        这是一种完全不同的类似数据库的方法,它使用水果和水果类型的查找表。这种方法可以处理任意数量的水果和水果类型。

        # create or read lookup table
        lut <- readr::read_table(
        "fruit    fruit_type
        banana           34
        apple            45
        orange           88")
        
        library(dplyr)
        library(tidyr)
        df %>% 
          mutate(fruit = fruits) %>% 
          separate_rows(fruit, sep = "\\s+") %>% 
          left_join(lut, by = "fruit") %>% 
          group_by(ID) %>% 
          mutate(rowid = row_number(ID)) %>% 
          pivot_wider(id_cols = c(ID, fruits), values_from = fruit_type, 
                      names_prefix = "fruit_type", names_from = rowid)
        
             ID fruits              fruit_type1 fruit_type2 fruit_type3
          <dbl> <chr>                     <dbl>       <dbl>       <dbl>
        1     1 banana apple orange          34          45          88
        2     2 apple orange                 45          88          NA
        3     3 orange                       88          NA          NA
        4     4 orange apple                 88          45          NA
        5     5 nothing                      NA          NA          NA
        

        fruits 列被复制然后拆分。现在,fruit 列在不同的行中包含一个水果。这些与查找表lut 连接以获得匹配的fruit_type 值。在将此结果重新调整为宽格式之前,需要对新列进行编号。这是通过对每个ID 中的行编号来实现的。

        编辑:

        根据OP's comment,生产数据集包含的段落中的关键字不是由空格分隔,而是由逗号等标点符号或以复数形式出现并带有尾随s。另外,关键词可以大写,也可以在一个段落中出现多次。

        我们可以尝试从段落中提取关键字,而不是将所有单词分开。这可以通过将所有关键字组合到一个 正则表达式中 | 来实现。因此,正则表达式 banana|apple|orange 将匹配其中一个结果。

        为了测试,我们需要一个更复杂的用例:

        df <- tibble(fruits = readr::read_lines(
        "There are bananas, oranges, and also apples here
        One Orange and another orange make two Oranges 
        apples and pineapples go together
        But pineapples alone must not be counted
        banana apple orange
        apple orange
        orange
        orange apple
        nothing")
        ) %>% 
          mutate(ID = row_number())
        

        加上修改后的代码

        df %>% 
          mutate(fruit = fruits %>% 
                   tolower() %>% 
                   stringr::str_extract_all(paste(lut$fruit, collapse = "|")) %>% 
                   lapply(unique)) %>% 
          unnest(fruit, keep_empty = TRUE) %>% 
          left_join(lut, by = "fruit") %>% 
          group_by(ID) %>% 
          mutate(rowid = row_number(ID)) %>% 
          pivot_wider(id_cols = c(ID, fruits), values_from = fruit_type, 
                      names_prefix = "fruit_type", names_from = rowid)
        

        我们得到

             ID fruits                                             fruit_type1 fruit_type2 fruit_type3
          <int> <chr>                                                    <dbl>       <dbl>       <dbl>
        1     1 "There are bananas, oranges, and also apples here"          34          88          45
        2     2 "One Orange and another Orange make two Oranges "           88          NA          NA
        3     3 "apples and pineapples go together"                         45          NA          NA
        4     4 "But pineapples alone must not be counted"                  45          NA          NA
        5     5 "banana apple orange"                                       34          45          88
        6     6 "apple orange"                                              45          88          NA
        7     7 "orange"                                                    88          NA          NA
        8     8 "orange apple"                                              88          45          NA
        9     9 "nothing"                                                   NA          NA          NA
        

        这种方法检测到了复数形式的关键字,与大小写无关。

        请注意,我特意选择了 lapply(unique) 仅计算一个段落中关键字的多次出现次数。如果每次出现都要单独计算,那么只需删除该行代码即可。

        但是,这种方法有一个(至少)缺点:单词pineapple 被视为apple,因为它包含apple 作为子字符串。

        【讨论】:

        • 这与我想要做的非常接近,但我遇到了separate_rows 函数的问题。在我的数据集中,每个“水果”关键字都是段落的一部分,所以我不能按每个单词分开。例如,像“这里有香蕉、橙子和苹果”这样的句子,所以我不想为每个单词创建一个新行;仅适用于那些特定的水果关键字。有什么方法可以用separate_rows 做到这一点?
        • @emv7 我明白了。问题不仅在于逗号和其他标点符号,还在于复数形式的尾随 s。
        • 谢谢,当我将它粘贴到 Rstudio 中时,我得到了不同的结果。它通过fruit_type15创建了fruit_type1,而不仅仅是fruit_type1、2和3。我很难发布完整的结果,但它在每列中只粘贴一个fruit_type编号。
        • @emv7 通过注释掉 `group_by(ID) %>%` 行,我能够创建类似的结果(15 列,每列只有一个数字)。
        【解决方案5】:

        您可以在 {tidyr} 中使用一些有用的功能,例如pivot_longer(),一旦您拥有一个宽数据框。 (还有一个函数pivot_wider()做相反的事情。)

        以下解决方案首先创建一个更宽的数据框,然后将其缩小为更长的数据框。因此它会按照您列出的相反顺序生成数据帧。

        library(dplyr)
        library(tidyr)
        
        ID <- c(1, 2, 3, 4, 5)
        fruits <- c("banana apple orange",
                    "apple orange",
                    "orange",
                    "orange apple",
                    "nothing")
        df <- tibble(ID, fruits)
        
        new_df <- 
          df %>%
          mutate(fruit_type = if_else(grepl("banana", fruits), 34, NA_real_),
                 fruit_type2 = if_else(grepl("apple", fruits), 45, NA_real_),
                 fruit_type3 = if_else(grepl("orange", fruits), 88, NA_real_))
        new_df
        #> # A tibble: 5 x 5
        #>      ID fruits              fruit_type fruit_type2 fruit_type3
        #>   <dbl> <chr>                    <dbl>       <dbl>       <dbl>
        #> 1     1 banana apple orange         34          45          88
        #> 2     2 apple orange                NA          45          88
        #> 3     3 orange                      NA          NA          88
        #> 4     4 orange apple                NA          45          88
        #> 5     5 nothing                     NA          NA          NA
        
        long_df <-
          new_df %>%
          pivot_longer(cols = starts_with("fruit_type"), names_to = "fruit_type") %>%
          select(-fruit_type) %>%
          rename(fruit_type = value) %>%
          distinct() %>%  # Remove duplicates
          group_by(ID, fruits) %>%
          mutate(n = n()) %>%
          filter(!is.na(fruit_type) | n == 1) %>%
          select(-n)
        long_df
        #> # A tibble: 9 x 3
        #> # Groups:   ID, fruits [5]
        #>      ID fruits              fruit_type
        #>   <dbl> <chr>                    <dbl>
        #> 1     1 banana apple orange         34
        #> 2     1 banana apple orange         45
        #> 3     1 banana apple orange         88
        #> 4     2 apple orange                45
        #> 5     2 apple orange                88
        #> 6     3 orange                      88
        #> 7     4 orange apple                45
        #> 8     4 orange apple                88
        #> 9     5 nothing                     NA
        

        由reprex package (v2.0.0) 于 2021-07-23 创建

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2018-05-31
          • 2017-11-08
          • 1970-01-01
          • 2020-11-23
          • 1970-01-01
          • 1970-01-01
          • 2020-06-15
          相关资源
          最近更新 更多