【问题标题】:R - How to replace a string from multiple matches (in a data frame)R - 如何从多个匹配项中替换字符串(在数据框中)
【发布时间】:2017-08-17 08:10:31
【问题描述】:

我需要用存储在数据框中的一些匹配项替换字符串的子集。

例如——

input_string = "Whats your name and Where're you from"

我需要从数据框中替换此字符串的一部分。说数据框是

matching <- data.frame(from_word=c("Whats your name", "name", "fro"),
            to_word=c("what is your name","names","froth"))

预期输出是你叫什么名字和你来自哪里

注意-

  1. 是匹配最大字符串。在此示例中,name 与 names 不匹配,因为 name 是更大匹配的一部分
  2. 它必须匹配整个字符串而不是部分字符串。 “from”的fro不应与“froth”匹配

我参考了以下链接,但不知何故无法按上述预期/描述完成这项工作

Match and replace multiple strings in a vector of text without looping in R

这是我在这里的第一篇文章。如果我没有提供足够的细节,请告诉我

【问题讨论】:

    标签: r string replace gsub


    【解决方案1】:
    toreplace =list("x1" = "y1","x2" = "y2", ..., "xn" = "yn")
    

    函数有两个参数 xi 和 yi。

    xi 是模式(查找什么),
    yi 是替换(替换为)。

    input_string = "Whats your name and Where're you from"
    toreplace<-list("Whats your name" = "what is your name", "names" = "name", "fro" = "froth")
    gsubfn(paste(names(toreplace),collapse="|"),toreplace,input_string)
    

    【讨论】:

    • 感谢 Aleksandr Voitov 的回复。我认为“from”中的“fro”也正在被替换。这就是我得到的答案,你叫什么名字,你在哪里frothm
    • 没问题@Sri。这是非常有用的功能,当您想用有意义的东西替换特定字符串时,它可以防止多次使用 gsub()。
    • 哦,抱歉,我的评论可能不太清楚。在您的代码中,最后一个 from 不应变为 froth。请问有什么办法可以防止
    • 在您的问题中,您提到要将“fro”替换为“froth”。那么真正的替代品应该是什么?
    • 对不起,我说过不应该匹配。预期输出是你叫什么名字和你来自哪里。另外,如果匹配的数据框运行到 200 万行,这会表现良好吗
    【解决方案2】:

    正在尝试不同的东西,下面的代码似乎有效。

    a <-c("Whats your name", "name", "fro")
    b <- c("what is your name","names","froth")
    c <- c("Whats your name and Where're you from")
    
    for(i in seq_along(a)) c <- gsub(paste0('\\<',a[i],'\\>'), gsub(" ","_",b[i]), c)
    c <- gsub("_"," ",c)
    c
    

    从下面的链接Making gsub only replace entire words?获得帮助

    但是,如果可能,我想避免循环。有人可以改进这个答案,没有循环

    【讨论】:

      【解决方案3】:

      编辑

      根据 Sri 的评论,我建议使用:

      library(gsubfn)
      # words to be replaced
      a <-c("Whats your","Whats your name", "name", "fro")
      # their replacements
      b <- c("What is yours","what is your name","names","froth")
      # named list as an input for gsubfn
      replacements <- setNames(as.list(b), a)
      # the test string
      input_string = "fro Whats your name and Where're name you from to and fro I Whats your"
      # match entire words
      gsubfn(paste(paste0("\\w*", names(replacements), "\\w*"), collapse = "|"), replacements, input_string)
      

      原创

      我不会说这比你的简单循环更容易阅读,但它可能会更好地处理重叠替换:

      # define the sample dataset
      input_string = "Whats your name and Where're you from"
      matching <- data.frame(from_word=c("Whats your name", "name", "fro", "Where're", "Whats"),
                             to_word=c("what is your name","names","froth", "where are", "Whatsup"))
      
      # load used library
      library(gsubfn)
      
      # make sure data is of class character
      matching$from_word <- as.character(matching$from_word)
      matching$to_word <- as.character(matching$to_word)
      
      # extract the words in the sentence
      test <- unlist(str_split(input_string, " "))
      # find where individual words from sentence match with the list of replaceble words
      test2 <- sapply(paste0("\\b", test, "\\b"), grepl, matching$from_word)
      # change rownames to see what is the format of output from the above sapply
      rownames(test2) <- matching$from_word
      # reorder the data so that largest replacement blocks are at the top
      test3 <- test2[order(rowSums(test2), decreasing = TRUE),]
      # where the word is already being replaced by larger chunk, do not replace again
      test3[apply(test3, 2, cumsum) > 1] <- FALSE
      
      # define the actual pairs of replacement
      replacements <- setNames(as.list(as.character(matching[,2])[order(rowSums(test2), decreasing = TRUE)][rowSums(test3) >= 1]),
                               as.character(matching[,1])[order(rowSums(test2), decreasing = TRUE)][rowSums(test3) >= 1])
      
      # perform the replacement
      gsubfn(paste(as.character(matching[,1])[order(rowSums(test2), decreasing = TRUE)][rowSums(test3) >= 1], collapse = "|"),
             replacements,input_string)
      

      【讨论】:

      • 谢谢@ira。我在您的代码中注意到了两点,1. 我使用具有 4500 多行的匹配数据框进行了测试。我的循环方式在 0.2 秒内执行,上面的代码用了 0.4 秒。和 2. 我认为您的代码期望字符串的顺序与 a 的顺序相同。例如,如果我将输入字符串指定为“to and fro is the name” - 您的代码会将 fro 替换为 froth,但不会将 name 替换为名称。我认为这是因为订购?我不确定。
      • @Sri 错误是因为在代码中,我没有费心去验证整个模式是否匹配,或者只是其中的一部分。但是我现在根据 Alekandr Voitov 的回答提出了更优雅的方法,这应该可以从他的回答中解决问题。
      • 谢谢@ira。有用。我必须补充一点,如果替换“from”和“to”的行数很少,它就可以工作。但是当我尝试替换有 60,000 行时,它会出错,无法编译正则表达式。所以现在,我将使用循环继续我的解决方案本身,直到有人使它变得更好(通过删除循环)
      • @Sri 我明白了...在这种情况下请记住,您的循环取决于 a 向量的顺序,最大的替换必须先出现,否则它也不会正确替换
      • 建议你使用边界匹配\\b而不是\\w*
      猜你喜欢
      • 1970-01-01
      • 2019-12-10
      • 2013-02-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多