【问题标题】:Replace multiple strings in data.frame by values from another data.frame用另一个 data.frame 中的值替换 data.frame 中的多个字符串
【发布时间】:2018-02-02 20:27:41
【问题描述】:

我正在尝试将字符串 data.frame 中出现的字符串替换为另一个字符串 data.frame 中的另一个字符串。

应替换子字符串的多个基本字符串

# base strings which I want to replace
base  <- data.frame(cmd = rep("this is my example <repl1> and here second <repl2> ...", nrow(repl1)))

替换字符串

# definition of replacement strings
repl1 <- data.frame(as.character(1:10))
repl2 <- data.frame(as.character(10:1))

我尝试使用 lapply 遍历 data.frame...

# what I have tried
lapply(base, function(x) {gsub("<repl1>", repl1, x)})

结果我比关注...

 [1] "this is my example c(1, 3, 4, 5, 6, 7, 8, 9, 10, 2) and here second <repl2> ..."
 [2] "this is my example c(1, 3, 4, 5, 6, 7, 8, 9, 10, 2) and here second <repl2> ..."
 [3] "this is my example c(1, 3, 4, 5, 6, 7, 8, 9, 10, 2) and here second <repl2> ..."

但我想实现...

 [1] "this is my example 1 and here second 10 ..."
 [2] "this is my example 2 and here second 9 ..."
 [3] "this is my example 3 and here second 8 ..."

感谢每个建议 :)

【问题讨论】:

    标签: r string dataframe replace


    【解决方案1】:

    我们可以在这里使用矢量化的regmatches 函数。这将删除所有循环:

    首先,由于您的替换位于不同的数据帧中,请将它们组合在一起:

    repl3 <- cbind(A=repl1,B=repl2)
    

    我们还有一个问题。您创建数据框的方式,字符在类factor。所以我会改变它:

    s <- as.character(base$cmd)
    

    从这里我们直接替换:

     regmatches(s,gregexpr("<repl1>|<repl2>",s))<- strsplit(do.call(paste,repl3)," ")
    s
     [1] "this is my example 1 and here second 10 ..."
     [2] "this is my example 2 and here second 9 ..." 
     [3] "this is my example 3 and here second 8 ..." 
     [4] "this is my example 4 and here second 7 ..." 
     [5] "this is my example 5 and here second 6 ..." 
     [6] "this is my example 6 and here second 5 ..." 
     [7] "this is my example 7 and here second 4 ..." 
     [8] "this is my example 8 and here second 3 ..." 
     [9] "this is my example 9 and here second 2 ..." 
    [10] "this is my example 10 and here second 1 ..."
    

    您的数据中需要使用许多代码,因为每次创建数据框时,您都忘记使用stringsAsFactors=F 选项。如果你这样做了,那么代码会很简单:

    v=as.character(base$cmd)
    repl4=data.frame(1:10,10:1,stringsAsFactors=F)
    regmatches(v,gregexpr("<repl1>|<repl2>",v))<-data.frame(t(repl4))
    v
     [1] "this is my example 1 and here second 10 ..."
     [2] "this is my example 2 and here second 9 ..." 
     [3] "this is my example 3 and here second 8 ..." 
     [4] "this is my example 4 and here second 7 ..." 
     [5] "this is my example 5 and here second 6 ..." 
     [6] "this is my example 6 and here second 5 ..." 
     [7] "this is my example 7 and here second 4 ..." 
     [8] "this is my example 8 and here second 3 ..." 
     [9] "this is my example 9 and here second 2 ..." 
    [10] "this is my example 10 and here second 1 ..."
    

    【讨论】:

    • 你好,当然更优雅的方式,谢谢。也感谢您向我指出有关 stringsAsFactors=F 的问题。在我的情况下,替换没有分解。
    【解决方案2】:

    您需要同时索引基本数据帧和 repl1 数据帧。您的代码将整个 repl1 数据帧传递给基本数据帧的每一行。

    试试这个:

    # definition of replacement strings
    repl1 <- data.frame(as.character(1:10))
    repl2 <- data.frame(as.character(10:1))
    
    # base strings which I want to replace
    base  <- data.frame(cmd = rep("this is my example <repl1> and here second <repl2> ...", nrow(repl1)))
    
    answer<-sapply(1:nrow(repl1), function(x) {gsub("<repl1>", repl1[x,1],  base[x,1])})
    

    现在重复 answer 和 repl2 数据框

    补充: 另一种方法是 stringr 库中的 str_replace 函数:

    library(stringr)
    answer<-str_replace(base[,1], "<repl1>", as.character(repl1[,1]))
    

    这很可能比 sapply 方法更快。

    【讨论】:

    • 谢谢,正是我想要的:)
    猜你喜欢
    • 1970-01-01
    • 2015-07-15
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-04-06
    • 1970-01-01
    相关资源
    最近更新 更多