【发布时间】:2014-08-07 02:44:03
【问题描述】:
给定一个字符串向量texts 和一个模式向量patterns,我想为每个文本找到任何匹配的模式。
对于小型数据集,这可以在 R 中使用 grepl 轻松完成:
patterns = c("some","pattern","a","horse")
texts = c("this is a text with some pattern", "this is another text with a pattern")
# for each x in patterns
lapply( patterns, function(x){
# match all texts against pattern x
res = grepl( x, texts, fixed=TRUE )
print(res)
# do something with the matches
# ...
})
此解决方案是正确的,但无法按比例放大。即使有较大的数据集(约 500 个文本和模式),这段代码也非常慢,在现代机器上每秒只能解决大约 100 个案例 - 考虑到这是一个粗略的字符串部分匹配,没有正则表达式(使用 @ 设置),这很荒谬987654325@)。即使使lapply 并行也不能解决问题。
有没有办法高效地重写这段代码?
谢谢, 木龙
【问题讨论】:
-
你的模式总是单字吗?您是否只是对
patterns的每个元素是否出现在texts的一个或多个元素中感兴趣(或者您是否需要知道它们出现在texts的哪些元素中)?
标签: string r performance string-matching