【问题标题】:Remove words in one column present in another column in R删除 R 中另一列中存在的一列中的单词
【发布时间】:2018-10-01 14:49:39
【问题描述】:

我有一个采用这种格式的数据框:

A <- c("John Smith", "Red Shirt", "Family values are better")
B <- c("John is a very highly smart guy", "We tried the tea but didn't enjoy it at all", "Family is very important as it gives you values")

df <- as.data.frame(A, B)

我的目的是将结果返回为:

ID   A                           B
1    John Smith                  is a very highly smart guy
2    Red Shirt                   We tried the tea but didn't enjoy it at all
3    Family values are better    is very important as it gives you

我试过了:

test<-df %>% filter(sapply(1:nrow(.), function(i) grepl(A[i], B[i])))

但它并没有给我想要的东西。

有什么建议/帮助吗?

【问题讨论】:

    标签: r regex dataframe


    【解决方案1】:

    一种解决方案是将mapply 与strsplit 一起使用。

    诀窍是将df$A 拆分为单独的单词,然后折叠由| 分隔的单词,然后将其用作gsub 中的pattern 以替换为""。

    lst <- strsplit(df$A, split = " ")
    
    df$B <- mapply(function(x,y){gsub(paste0(x,collapse = "|"), "",df$B[y])},lst,1:length(lst))
    df
    # A                                           B
    # 1               John Smith                  is a very highly smart guy
    # 2                Red Shirt We tried the tea but didn't enjoy it at all
    # 3 Family values are better          is very important as it gives you 
    

    另一种选择是:

    df$B <- mapply(function(x,y)gsub(x,"",y) ,gsub(" ", "|",df$A),df$B)
    

    数据:

    A <- c("John Smith", "Red Shirt", "Family values are better")
    B <- c("John is a very highly smart guy", "We tried the tea but didn't enjoy it at all", "Family is very important as it gives you values")
    
    df <- data.frame(A, B, stringsAsFactors = FALSE)
    

    【讨论】:

    • 很好的答案@MKR,在我的情况下效果很好。在我的情况下,您的第二个选项工作得稍微快一些。我正在使用几百万行 data.frame。如果基于 data.table 的方法更省时,有什么线索吗?
    • 通过切换$A 和$B 这对我有用。谢谢!如何仅删除不匹配的部分?
    【解决方案2】:

    使用stringr::str_split_fixed 函数的另一种选择:

    library(stringr)
    
    str_split_fixed(sapply(paste(df$A,df$B, sep=" columnbreaker "), 
                    function(i){
                                paste(unique(
                                             strsplit(as.character(i), split=" ")[[1]]), 
                             collapse = " ")}), 
                     " columnbreaker ", 2)
    
    
    #       [,1]                       [,2]                                         
    # [1,] "John Smith"               "is a very highly smart guy"                 
    # [2,] "Red Shirt"                "We tried the tea but didn't enjoy it at all"
    # [3,] "Family values are better" "is very important as it gives you"  
    

    【讨论】:

    • 这也成功了。谢谢你。虽然我觉得上面的解决方案更优雅,更适合我的需要(我有一个巨大的数据集)
    猜你喜欢
    • 2021-11-18
    • 2019-08-31
    • 2021-02-03
    • 1970-01-01
    • 1970-01-01
    • 2022-06-23
    • 1970-01-01
    • 1970-01-01
    • 2021-10-23
    相关资源
    最近更新 更多