【问题标题】:Fast way for String matching and replacement from another dataframe in R从 R 中的另一个数据帧进行字符串匹配和替换的快速方法
【发布时间】:2018-05-09 16:51:55
【问题描述】:

我有两个看起来像这样的数据帧(虽然第一个数据帧超过 9000 万行,第二个数据帧超过 1400 万行)另外第二个数据帧是随机排序的

df1 <- data.frame(
  datalist = c("wiki/anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/individualism to complete wiki/collectivism",
               "strains of anarchism have often been divided into the categories of wiki/social_anarchism and wiki/individualist_anarchism or similar dual classifications",
               "the word is composed from the word wiki/anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e",
               "anarchy from anarchos meaning one without rulers from the wiki/privative prefix wiki/privative_alpha an- i.e",
               "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/infinitive suffix -izein",
               "the first known use of this word was in 1539"),
  words = c("anarchist_schools_of_thought  individualism  collectivism", "social_anarchism  individualist_anarchism",
            "anarchy  -ism", "privative  privative_alpha", "infinitive", ""),

  stringsAsFactors=FALSE)

df2 <- data.frame(
  vocabword = c("anarchist_schools_of_thought", "individualism","collectivism" , "1965-66_nhl_season_by_team","social_anarchism","individualist_anarchism",                
                 "anarchy","-ism","privative","privative_alpha", "1310_the_ticket",  "infinitive"),
  token = c("Anarchist_schools_of_thought" ,"Individualism", "Collectivism",  "1965-66_NHL_season_by_team", "Social_anarchism", "Individualist_anarchism" ,"Anarchy",
            "-ism", "Privative" ,"Alpha_privative", "KTCK_(AM)" ,"Infinitive"), 
  stringsAsFactors = F)

我能够将短语“wiki/”之后的所有单词提取到另一列中。这些单词需要替换为与第二个数据框中的 vocabword 匹配的标记列。因此,例如,我会查看第一个数据帧第一行中 wiki/ 之后的作品“anarchist_schools_of_thought”,然后在第二个数据帧中的词汇单词下找到术语“anarchist_schools_of_thought”,我想用相应的替换它令牌是“Anarchist_schools_of_thought”。

所以它最终应该是这样的:

1 wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism
2 strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications
3 the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e
4 anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative an- i.e
5 authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein
6 the first known use of this word was in 1539

我意识到其中很多只是将单词的第一个字母大写,但其中一些明显不同。我可以做一个 for 循环,但我认为这会花费太多时间,我更喜欢使用 data.table 方式或可能是 stringi 或 stringr 方式。而且我通常只会进行合并,但是由于需要在一行中替换多个单词,这会使事情变得复杂。

提前感谢您的帮助。

【问题讨论】:

  • 这与您昨天提出的问题有何不同? stackoverflow.com/q/50241313/5325862
  • 我需要替换文本。我想如果我把一些文字分开的话我可以弄明白,但我一直在努力,但一无所获。
  • 你能把上一篇文章中的代码加上你从那以后做了什么吗?这样我们就知道您在哪里停下来以及下一步需要什么
  • 我用这个:data$words = trimws(gsub("wiki/(\\S+)|(?:(?!wiki/\\S).)+", " \\1 ", data$datalist, perl=TRUE)) 就是这样。就像我说的,我可以做一个 for 循环,但它会非常慢。从昨天开始,我真的没有任何进展
  • 这很好,但该代码对于包含在问题正文中很重要。循环肯定会比向量操作慢,所以我认为上一篇文章让你走在了正确的轨道上

标签: r stringi


【解决方案1】:

您可以使用来自stringr 的str_replace_all 执行此操作:

library(stringr)

str_replace_all(df1$datalist, setNames(df2$vocabword, df2$token))

基本上,str_replace_all 允许您提供一个命名向量,其中原始字符串是名称,替换是向量的元素。您通过创建字符串和替换的“字典”完成了所有艰苦的工作。 str_replace_all 简单地接受了它并自动进行替换。

结果:

[1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism"              
[2] "strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications"
[3] "the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e"                               
[4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Privative_alpha an- i.e"                                              
[5] "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein"                                       
[6] "the first known use of this word was in 1539"

【讨论】:

  • 有字符串变体吗?我已经在 1000 行上运行了一分钟,它还在运行
  • @Kayla 我注意到“1965-66_nhl_season_by_team”从未出现在您的datalist 中,您是否故意将其添加为非匹配项?
  • 是的,我补充说不匹配
  • 如果所有的 df2 都是完全随机的,有没有办法做到这一点?也有 1400 万行
  • @Kayla 如果不匹配项没有出现在df1$words 中,逻辑是否可以自动排除不匹配项?
【解决方案2】:

这个问题的解决方案似乎适用于您的数据:R: replacing multiple regex with sub

install.packages('qdap')
qdap::mgsub(df2[,1], df2[,2], df1[,1])

[1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism"              
[2] "strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications"
[3] "the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e"                               
[4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_Privative an- i.e"                                              
[5] "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein"                                       
[6] "the first known use of this word was in 1539"            

【讨论】:

    【解决方案3】:

    因为您的每个术语都以“wiki/”开头,所以可以重新排列您的数据集,以便更轻松地创建匹配项。我提出的方法是将每个“wiki/term”移动到它自己的数据框行,使用连接来匹配有效的单词,然后反转步骤以将字符串重新组合在一起但是里面有新的术语。

    library(tidyverse)
    df1a <- df1 %>%
      # Create a separator character to identify where to split
      mutate(datalist = str_replace_all(datalist,"wiki/","|wiki/")) %>% 
      mutate(datalist = str_remove(datalist,"^\\|"))
    
      # Split so that each instance gets its own column
    df1a <- 
      str_split(df1a$datalist,"\\|",simplify = TRUE) %>% 
      as.tibble() %>% 
      # Add a rownum column to keep track where to put back together for later
      mutate(rownum = 1:n()) %>% 
      # Gather the dataframe into a tidy form to prepare for joining
      gather("instance","text",-rownum,na.rm = TRUE) %>% 
      # Create a column for joining to the data lookup table
      mutate(keyword = text %>% str_extract("wiki/[^ ]+") %>% str_remove("wiki/")) %>% 
      # Join the keywords efficiently using left_bind
      left_join(df2,by = c("keyword" = "vocabword")) %>% 
      # Put the results back into the text string
      mutate(text = str_replace(text,"wiki/[^ ]+",paste0("wiki/",token))) %>%
      select(-token,-keyword) %>% 
      # Spread the data back out to the original number of rows
      spread(instance,text) %>% 
      # Re-combine the sentences/strings to their original form
      unite("datalist",starts_with("V"),sep="") %>%
      select("datalist")
    

    结果:

    # A tibble: 6 x 1
      datalist                                                                                                 
      <chr>                                                                                                    
    1 wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individ~
    2 strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Indiv~
    3 the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively~
    4 anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative an-~
    5 authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive su~
    6 the first known use of this word was in 1539    
    

    【讨论】:

    • df1 (6*170k) 中的 1,020,000 次观察和 df2 中的 12 次观察(无变化)导致我的笔记本电脑上的运行时间为 3.29 分钟。显然,向 df2 添加项会增加它使用的时间。但唯一增加时间的是 left_join 语句,它相当有效。如果 df2 的大小增加,其他代码行不会改变所需的时间。
    【解决方案4】:

    我通常使用直接stringi 的方式如下:

    library(stringi)
    
    Old <- df2[["vocabword"]]
    New <- df2[["token"]]
    
    stringi::stri_replace_all_regex(df1[["datalist"]],
                                    "\\b"%s+%Old%s+%"\\b",
                                    New,
                                    vectorize_all = FALSE)
    
    #[1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism"              
    #[2] "strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications"
    #[3] "the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e"                               
    #[4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative an- i.e"                                              
    #[5] "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein"                                       
    #[6] "the first known use of this word was in 1539"  
    

    理论上,您应该能够通过以正确的方式并行化来获得合理的改进,但您将无法比Nx 加速(其中N = #可用内核)。 -- 我的直觉是,将运行时间从大约 8 个月缩短到 15 天在实际意义上仍然没有真正帮助你。

    但是,如果您有 1400 万个潜在替换来生成超过 9000 万行,那么似乎可能需要一种根本不同的方法。任何句子中的最大单词数是多少?


    更新:添加一些示例代码来对潜在解决方案进行基准测试:

    使用stringi::stri_rand_lipsum() 添加额外的句子并使用stringi::stri_rand_strings() 添加额外的替换对可以更容易地看到增加的语料库大小和词汇量对运行时的影响。

    1000 句:

    • 1000 对替换:3.9 秒。
    • 10,000 对替换:36.5 秒。
    • 100,000 对替换:365.4 秒。

    我不会尝试 1400 万,但这应该有助于您评估替代方法是否可以扩展。

    library(stringi)
    
    ExtraSentenceCount <- 1e3
    ExtraVocabCount <- 1e4
    
    Sentences <- c("wiki/anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/individualism to complete wiki/collectivism",
                   "strains of anarchism have often been divided into the categories of wiki/social_anarchism and wiki/individualist_anarchism or similar dual classifications",
                   "the word is composed from the word wiki/anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e",
                   "anarchy from anarchos meaning one without rulers from the wiki/privative prefix wiki/privative_alpha an- i.e",
                   "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/infinitive suffix -izein",
                   "the first known use of this word was in 1539",
                   stringi::stri_rand_lipsum(ExtraSentenceCount))
    
    vocabword <- c("anarchist_schools_of_thought", "individualism","collectivism" , "1965-66_nhl_season_by_team","social_anarchism","individualist_anarchism",                
               "anarchy","-ism","privative","privative_alpha", "1310_the_ticket",  "infinitive",
               "a",
               stringi::stri_rand_strings(ExtraVocabCount,
                                          length = sample.int(8, ExtraVocabCount, replace = TRUE),
                                          pattern = "[a-z]"))
    
    token <- c("Anarchist_schools_of_thought" ,"Individualism", "Collectivism",  "1965-66_NHL_season_by_team", "Social_anarchism", "Individualist_anarchism" ,"Anarchy",
               "-ism", "Privative" ,"Alpha_privative", "KTCK_(AM)" ,"Infinitive",
               "XXXX",
               stringi::stri_rand_strings(ExtraVocabCount,
                                          length = 3,
                                          pattern = "[0-9]"))
    
    system.time({
      Cleaned <- stringi::stri_replace_all_regex(Sentences, "\\b"%s+%vocabword%s+%"\\b", token, vectorize_all = FALSE)
    })
    
    #   user  system elapsed 
    # 36.652   0.070  36.768 
    
    head(Cleaned)
    
    # [1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism 749 complete wiki/Collectivism"                
    # [2] "strains 454 anarchism have often been divided into the categories 454 wiki/Social_anarchism and wiki/Individualist_anarchism 094 similar dual classifications"
    # [3] "the word 412 composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek 190.546"                             
    # [4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative 358- 190.546"                                            
    # [5] "authority sovereignty realm magistracy and the suffix 094 -ismos -isma from the verbal wiki/Infinitive suffix -izein"                                         
    # [6] "the first known use 454 this word was 201 1539" 
    

    更新 2:下面的方法没有考虑到标签是另一个子字符串的可能性——即wiki/Individualist 和wiki/Individualist_anarchism 可能会给你错误的结果。我真正知道避免这种情况的唯一方法是使用正则表达式/单词替换前后跟单词边界 (\\b) 的完整单词,这不能基于固定字符串。

    一个可能会给您带来希望的选项依赖于这样一个事实,即您基本上将所有想要的替换都标有前缀wiki/。如果您的实际使用是这种情况,那么我们可以利用它并使用固定替换而不是正则表达式替换前后跟单词边界的完整单词(\\b)。 (这是必要的,以避免像“ism”这样的词汇出现在较长单词的一部分时被替换)

    使用与上面相同的列表:

    prefixed_vocabword <- paste0("wiki/",vocabword)
    prefixed_token <- paste0("wiki/",token)
    
    system.time({
      Cleaned <- stringi::stri_replace_all_fixed(Sentences, prefixed_vocabword, prefixed_token, vectorize_all = FALSE)
    })
    

    这将运行时间缩短到 10.4 秒,包含 1,000 个句子和 10,000 个替换,但由于运行时间仍在线性增长,因此您的数据大小仍需要数小时。

    【讨论】:

      猜你喜欢
      • 2016-12-10
      • 2021-08-11
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-11-21
      • 2015-08-17
      相关资源
      最近更新 更多