【问题标题】:Check multiple conditions in a string and return binary indicator [duplicate]检查字符串中的多个条件并返回二进制指示符[重复]
【发布时间】:2020-10-31 22:10:21
【问题描述】:

我想根据comment_1和comment_2中的字符对我的df的数据进行过滤和分组,并给出指标0或1作为最终结果。但是,有一些规则与过滤器一起出现。规则是:

  1. 如果行的comment_1 由apple 组成,而行的comment_2 由apple 组成,则为 好吧,然后使用price_1 减去price_2。如果减法后的数大于 20,则 result 将是 1,如果小于 20,则 result 将是 0

  2. 如果行的comment_1 由orange 组成并且comment_2 由apple 组成/ comment_1 由apple 组成,comment_2 由orange 组成,然后也使用price_1 减号 price_2。如果 sbutraction 后的数字大于 10,则 result 将是 1 否则 result 将是 0。

请注意,Apple 或 apple、Orange 或 Orange 无关紧要,因此代码也应考虑大写字母。

例如:

  1. 数据的第一行是apple (comment_1)到Apple (comment_2),price_1减去price_2的结果是13小于20,所以result会显示为0。

  2. 第二行数据是橙色的(comment_1) to Apple (comment_2),减去price_1和price_2后的结果是11,因为11大于10,所以最终的result会显示为1.

  3. 由于第4行price_1-price_2= 2小于10,所以结果为0。

我将我的 df 附在下面,最后的列结果是最终答案。

price_1 <- c(25, 33, 54, 24)
price_2 <- c(12, 22, 11, 22)
itemid <- c(22203, 44412,55364, 552115)
itembds <- as.integer(c("", 21344, "", ""))
comment_1 <- c("The apple is expensive", "The orange is sweet", "The Apple is nice", "the apple is not nice")
comment_2 <- c("23 The Apple was beside me", "The Apple was left behind", "The apple had rotten", "the Orange should be fine" )
result <- c(0, 1, 1, 0)

df <- data.frame(price_1, price_2, itemid, itembds, comment_1, comment_2, result)

【问题讨论】:

    标签: r if-statement conditional-statements


    【解决方案1】:

    这是一个简单的 if-else 语句。这是使用dplyr 和stringr 的解决方案。

    library(dplyr)
    library(stringr)
    
    df %>% mutate(price_a = if_else(str_detect(comment_1, "[Aa]pple") & 
                                  str_detect(comment_2, "[Aa]pple"), price_1 - price_2, 0),
                  price_o = if_else(str_detect(comment_1, "[Oo]range") &
                                      str_detect(comment_2, "[Aa]pple"), price_1 - price_2, 0),
                  price_o = if_else(str_detect(comment_1, "[Aa]pple") &
                                      str_detect(comment_2, "[Oo]range"), price_1 - price_2, price_o),
                  res_actual = if_else(price_o > 10, 1, 0),
                  res_actual = if_else(price_a > 20, 1, res_actual)) %>% 
      select(-price_o, -price_a)
    
    price_1 price_2 itemid itembds              comment_1                  comment_2 result res_actual
    1      25      12  22203      NA The apple is expensive 23 The Apple was beside me      0          0
    2      33      22  44412   21344    The orange is sweet  The Apple was left behind      1          1
    3      54      11  55364      NA      The Apple is nice       The apple had rotten      1          1
    4      24      22 552115      NA  the apple is not nice  the Orange should be fine      0          0
    

    【讨论】:

      【解决方案2】:

      使用case_when 和grepl 来测试各种条件。

      library(dplyr)
      
      df %>%
          mutate(result = case_when(
                   grepl('apple', comment_1, ignore.case = TRUE) &
                   grepl('apple', comment_2, ignore.case = TRUE) ~ +(price_1 - price_2 > 20), 
                   grepl('orange', comment_1, ignore.case = TRUE) &
                   grepl('apple', comment_2, ignore.case = TRUE) |
                   grepl('apple', comment_1, ignore.case = TRUE) &
                   grepl('orange', comment_2, ignore.case = TRUE) ~ +(price_1 - price_2 > 10)))
      
      
      #  price_1 price_2 itemid itembds              comment_1                  comment_2 result
      #1      25      12  22203      NA The apple is expensive 23 The Apple was beside me      0
      #2      33      22  44412   21344    The orange is sweet  The Apple was left behind      1
      #3      54      11  55364      NA      The Apple is nice       The apple had rotten      1
      #4      24      22 552115      NA  the apple is not nice  the Orange should be fine      0
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2021-06-18
        • 2013-02-11
        • 2019-02-02
        • 2021-03-27
        • 1970-01-01
        • 1970-01-01
        • 2023-03-02
        相关资源
        最近更新 更多