【问题标题】:String Matching with very large file in RR中非常大的文件的字符串匹配
【发布时间】:2019-02-28 02:08:46
【问题描述】:

我有一个非常大的 RDS 文章文件 (13GB)。 R 全局环境中的数据帧大小约为 6GB

每篇文章都有一个 ID、一个日期、带有 POS 标记的正文,一个只有两三个带有 POS 标记的单词的模式。和其他一些元数据。

structure(list(an = c("1", "2", "3", "4", "5"), pub_date = structure(c(11166, 8906, 12243, 4263, 13077), class = "Date"), 
source_code = c("1", "2", "2", "3", "2"), word_count = c(99L, 
97L, 30L, 68L, 44L), POStagged = c("the_DT investment_NN firm_NN lehman_NN brothers_NNS holdings_NNS said_VBD yesterday_NN that_IN it_PRP would_MD begin_VB processing_VBG its_PRP$ own_JJ stock_NN trades_NNS by_IN early_RB next_JJ year_NN and_CC end_VB its_PRP$ existing_VBG tradeclearing_NN contract_NN with_IN the_DT bear_NN stearns_VBZ companies_NNS lehman_NN which_WDT is_VBZ the_DT last_JJ big_JJ securities_NNS firm_NN to_TO farm_VB out_RP its_PRP$ stock_NN trade_NN processing_NN said_VBD it_PRP would_MD save_VB million_CD to_TO million_CD annually_RB by_IN clearing_VBG its_PRP$ own_JJ trades_NNS a_DT bear_NN stearns_VBZ spokesman_NN said_VBD lehmans_NNS business_NN contributed_VBD less_JJR than_IN percent_NN to_TO bear_VB stearnss_NN clearing_NN operations_NNS", 
"six_CD days_NNS after_IN she_PRP was_VBD introduced_VBN as_IN womens_NNS basketball_NN coach_NN at_IN wisconsin_NN with_IN a_DT fouryear_JJ contract_NN nell_NN fortner_NN resigned_VBD saying_VBG she_PRP wants_VBZ to_TO return_VB to_TO louisiana_JJR tech_NN as_IN an_DT assistant_NN im_NN shocked_VBN said_VBD associate_JJ athletic_JJ director_NN cheryl_NN marra_NN east_JJ carolina_NN came_VBD from_IN behind_IN with_IN two_CD runs_NNS in_IN the_DT seventh_JJ inning_NN and_CC defeated_VBD george_NN mason_NN in_IN the_DT colonial_JJ athletic_JJ association_NN baseball_NN tournament_NN in_IN norfolk_NN johnny_NN beck_NN went_VBD the_DT distance_NN for_IN the_DT pirates_NNS boosting_VBG his_PRP$ record_NN to_TO the_DT patriots_NNS season_NN closed_VBD at_IN", 
"tomorrow_NN clouds_NNS and_CC sun_NN high_JJ low_JJ", "the_DT diversity_NN of_IN the_DT chicago_NN financial_JJ future_NN markets_NNS the_DT chicagoans_NNS say_VBP also_RB enhances_VBG their_PRP$ strength_NN traders_NNS and_CC arbitragers_NNS can_MD exploit_VB price_NN anomalies_NNS for_IN example_NN between_IN cd_NN and_CC treasurybill_NN futures_NNS still_RB nyfe_JJ supporters_NNS say_VBP their_PRP$ head_NN start_VB in_IN cd_NN futures_NNS and_CC technical_JJ advantages_NNS in_IN the_DT contract_NN traded_VBN on_IN the_DT nyfe_NN mean_VBP that_IN the_DT chicago_NN exchanges_NNS will_MD continue_VB to_TO play_VB catchup_NN", 
"williams_NNS industries_NNS inc_IN the_DT manufacturing_NN and_CC construction_NN company_NN provides_VBZ steel_NN products_NNS to_TO build_VB major_JJ infrastructure_NN it_PRP has_VBZ been_VBN involved_VBN with_IN area_NN landmark_NN projects_NNS including_VBG rfk_JJ stadium_NN left_VBD the_DT woodrow_JJ wilson_NN bridge_NN and_CC the_DT mixing_NN bowl_NN"
), phrases = c("begin processing", "wants to return", "high", 
"head start in", "major"), repeatPhraseCount = c(1L, 1L, 
1L, 1L, 1L), pattern = c("begin_V", "turn_V", "high_JJ", 
"start_V", "major_JJ"), code = c(NA_character_, NA_character_, 
NA_character_, NA_character_, NA_character_), match = c(TRUE, 
TRUE, TRUE, TRUE, TRUE)), .Names = c("an", "pub_date", "source_code", "word_count", "POStagged", "phrases", "repeatPhraseCount", "pattern", 
"code", "match"), row.names = c("4864065", "827626", "6281115", 
"281713", "3857705"), class = "data.frame")

我的目标是(针对每一行)检测 POSTtag 中是否存在模式。

模式列是我个人构建的固定列表。该列表是 465 个单词/短语及其 POS。

我想进行匹配,以便在将 doubt 等词用作动词或名词时进行区分。基本上是为了确定上下文。

但是,在某些情况下,我有短语而不是单词,短语的结尾可能是一个不断变化的模式。例如,短语“可能无法达成交易”,其中“能够达成交易”可以是任何动词短语(例如 be能够达成交易)。我的尝试多种多样,不确定我是否以正确的方式进行:

--might_MD not_RB _VP (this works and picks up ***might not*** but is clearly wrong since the verb phrase after it is not picked)

如果我使用 fixed() ,那么 str_detect 就可以工作并且执行速度非常快。但是,fixed() 肯定会丢失某些情况(如上所述),我无法确定比较结果。这是一个例子:

str_detect("might_MD not_RB be able to make the deal", "might_MD not_RB [A-Za-z]+(?:\\s+[A-Za-z]+){0,6}")
TRUE

str_detect("might_MD not_RB be able to make the deal", fixed("might_MD not_RB [A-Za-z]+(?:\\s+[A-Za-z]+){0,6}"))
FALSE

https://stackoverflow.com/a/51406046/3290154

我想要的输出是我的数据框中的一个附加列,其 TRUE/FALSE 结果告诉我是否在 POSTtag 中看到了模式。

## Attempt 1 - R fatally crashes
## this works in a smaller sample but bombs R in a large dataframe
df$match <- str_detect(df$POStagged, df$pattern)

## Attempt 2
## This bombs (using multidplyr and skipping some lines of code)
partition(source_code, cluster=cl) %>%
    mutate(match=str_detect(POStagged, pattern)) %>%
    filter(!(match==FALSE)) %>%
    filter(!is.na(match)) %>%
    collect()

##I get this error: Error in serialize(data, node$con) : error writing to connection

根据我的理解,这是由于 multidplyr 处理内存的方式以及它如何在内存中加载数据的限制 (https://github.com/hadley/multidplyr/blob/master/vignettes/multidplyr.md)。但是,由于 multidplyr 使用的是并行包,如果我在这里推断,我应该仍然可以 - 如果我将我的数据分成 5 个副本,那么 6*5 = 30GB 加上任何包等等。

## Attempt 3 - I tried to save the RDS to a csv/txt file and use the chuncked package, however, the resulting csv/txt was over 100GB.

## Attempt 4 - I tried to run a for loop, but I estimate it will take ~12days to run

我阅读了一些关于正则表达式的贪婪的内容,因此我尝试通过附加 ?+ 来修改我的模式列(使我的正则表达式变得懒惰)。但是,走这条路意味着我不能使用 fixed() ,因为我所有的匹配都是错误的。非常感谢您在正确方向上的任何帮助!

https://stringr.tidyverse.org/articles/regular-expressions.html

What do 'lazy' and 'greedy' mean in the context of regular expressions?

【问题讨论】:

  • 我正在尝试根据您的代码了解您的目标,但我不确定我是否明白。请用文字说明一下好吗?似乎您正在尝试检测和标记数据框中的所有行,其中(一些?全部?)pattern 列中的空格分隔字符串出现在POStagged 列中。它是否正确?而您正在使用str_detect... 因为您认为它会比grepl 更快?如果您共享几行数据(例如 5-10 行)并获得所需的结果,这也会有所帮助。没有看到这一点,很难确定fixed() 是否是一个可行的选择。
  • 当你似乎只给它一个字符串列作为输入时,你为什么在preprocess 中使用lapply?我不确定您正在运行它,因为您在 df$variable 上运行它,但您的示例数据不包含名为 variable... 的列 df$variable 是列表列吗?否则lapply 似乎是一个巨大的低效率。当您共享更多示例数据时,请以明确列类的方式进行 - dput() 最适合此操作,因为它提供了确切数据结构的复制/可粘贴版本。
  • 谢谢@Gregor - 我已经包含了更多信息
  • 这个新例子很有帮助。仍然存在一些问题:(1)我不知道您所说的 “我不想要完全匹配,因此例如,我想检测“可能”以及“非常可能”是什么意思 i>. 您的数据中既没有出现“可能”也没有出现“非常可能” - 这应该是要匹配的字符串的示例,还是您对匹配实际匹配的可能性有多大含糊?需要匹配吗?你能举出你仍然想捕捉的非精确匹配的例子吗?
  • (2) 您示例中的前三个模式看起来像是单个术语(我认为?),但第四个模式是"the_DT _JJS NP"。您是否需要找到整个术语,或者您正在寻找,在任何地方说出所有“the_DT` 和_JJS 和NP,但不一定是连续的?(这就是patternList,它出现在某些地方你的代码而不是你的数据在做什么?)

标签: regex performance string-matching stringr regex-greedy


【解决方案1】:

也许你在使用的时候可以取得更快的进步,得到更好的结果 Open Sourcing BERT: State-of-the-Art Pre-training for Natural Language Processing 而不是?这是一种完全不同的方法,当然,我知道,对不起。因此,以防万一您由于某种原因不知道它。

【讨论】:

    猜你喜欢
    • 2011-02-09
    • 1970-01-01
    • 1970-01-01
    • 2020-10-23
    • 2015-03-23
    • 2016-10-08
    • 1970-01-01
    • 2019-10-06
    • 2022-01-02
    相关资源
    最近更新 更多