【问题标题】:Test two columns of strings for match row-wise in R在 R 中测试两列字符串以逐行匹配
【发布时间】:2015-06-15 20:29:51
【问题描述】:

假设我有两列字符串:

library(data.table)
DT <- data.table(x = c("a","aa","bb"), y = c("b","a","bbb"))

对于每一行,我想知道 x 中的字符串是否存在于 y 列中。循环方法是:

for (i in 1:length(DT$x)){
  DT$test[i] <- DT[i,grepl(x,y) + 0]
}

DT
    x   y test
1:  a   b    0
2: aa   a    0
3: bb bbb    1

这个有矢量化的实现吗?使用grep(DT$x,DT$y) 只使用x 的第一个元素。

【问题讨论】:

  • @LegalizeIt 是的,我不是在寻找 a == b 的情况,而是在行之间存在部分字符串匹配的情况。以上测试 x 是否在 y 中,但如果这很重要,您可以重新运行其他方式。我不确定您所说的“为什么 y 列中没有“a”?”
  • @DavidArenburg 好主意,但仅适用于 x 唯一时。我试过DT[, test := grepl(x, y) + 0, by = .I] 但得到相同的argument 'pattern' has length &gt; 1 and only the first element will be used 错误。也就是说,可能有一个解决方案,您首先调用 DT[, RowI := .I],然后使用 by = rowI
  • 它不仅在x 唯一时有效。例如,对于DT &lt;- data.table(x = c("a","aa","aa","bb"), y = c("b","a","a", "bbb")) 非常有效。
  • @DavidArenburg 你是对的......它会为每组 x 将相同的参数传递给 grepl,这都是相同的输入。聪明的。我正在对所有答案进行基准测试,并将给出最好的答案。
  • 哦,我知道你在那里做了什么。你问这个赏金问题......

标签: regex r data.table


【解决方案1】:

你可以这样做

DT[, test := grepl(x, y), by = x]

【讨论】:

    【解决方案2】:

    或者mapply(Vectorize实际上只是mapply的包装)

    DT$test <- mapply(grepl, pattern=DT$x, x=DT$y)
    

    【讨论】:

      【解决方案3】:

      感谢大家的回复。我已经对它们进行了基准测试,并得出以下结论:

      library(data.table)
      library(microbenchmark)
      
      DT <- data.table(x = rep(c("a","aa","bb"),1000), y = rep(c("b","a","bbb"),1000))
      
      DT1 <- copy(DT)
      DT2 <- copy(DT)
      DT3 <- copy(DT)
      DT4 <- copy(DT)
      
      microbenchmark(
      DT1[, test := grepl(x, y), by = x]
      ,
      DT2$test <- apply(DT, 1, function(x) grepl(x[1], x[2]))
      ,
      DT3$test <- mapply(grepl, pattern=DT3$x, x=DT3$y)
      ,
      {vgrepl <- Vectorize(grepl)
      DT4[, test := as.integer(vgrepl(x, y))]}
      )
      

      结果

      Unit: microseconds
                                                                                     expr       min        lq       mean     median        uq        max neval
                                                   DT1[, `:=`(test, grepl(x, y)), by = x]   758.339   908.106   982.1417   959.6115  1035.446   1883.872   100
                                  DT2$test <- apply(DT, 1, function(x) grepl(x[1], x[2])) 16840.818 18032.683 18994.0858 18723.7410 19578.060  23730.106   100
                                    DT3$test <- mapply(grepl, pattern = DT3$x, x = DT3$y) 14339.632 15068.320 16907.0582 15460.6040 15892.040 117110.286   100
       {     vgrepl <- Vectorize(grepl)     DT4[, `:=`(test, as.integer(vgrepl(x, y)))] } 14282.233 15170.003 16247.6799 15544.4205 16306.560  26648.284   100
      

      除了语法上最简单之外,data.table 解决方案也是最快的。

      【讨论】:

      • 我不确定这里的礼仪是什么(接受大卫的回答或我的回答)。无论如何,为所有人 +1 表示可以完成此操作的方法的数量。
      • 对于这个特定的例子(x 和 y 值都重复),我看到在第一个例子中用 by=.(x,y) 代替 by=x 略有改进。
      【解决方案4】:

      您可以将grepl 函数传递给应用函数,以对数据表的每一行进行操作,其中第一列包含要搜索的字符串,第二列包含要搜索的字符串。这应该会给您一个问题的矢量化解决方案。

      > DT$test <- apply(DT, 1, function(x) as.integer(grepl(x[1], x[2])))
      > DT
          x   y test
      1:  a   b    0
      2: aa   a    0
      3: bb bbb    1
      

      【讨论】:

        【解决方案5】:

        你可以使用Vectorize:

        vgrepl <- Vectorize(grepl)
        DT[, test := as.integer(vgrepl(x, y))]
        DT
            x   y test
        1:  a   b    0
        2: aa   a    0
        3: bb bbb    1
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2019-03-12
          • 2013-05-29
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多