【问题标题】:Tokenization in r tidytext, leaving in ampersandsr tidytext 中的标记化,留下 & 符号
【发布时间】:2020-08-04 16:48:20
【问题描述】:

我目前正在使用tidytext 包中的unnest_tokens() 函数。它完全按照我的需要工作,但是,它从文本中删除了与号 (&)。我希望它不要那样做,但保持其他一切不变。

例如:

library(tidyverse)
library(tidytext)

d <- tibble(txt = "Let's go to the Q&A about B&B, it's great!")
d %>% unnest_tokens(word, txt, token="words")

目前返回

# A tibble: 11 x 1
   word 
   <chr>
 1 let's
 2 go   
 3 to   
 4 the  
 5 q    
 6 a    
 7 about
 8 b    
 9 b    
10 it's 
11 great

但我希望它返回

# A tibble: 9 x 1
  word 
  <chr>
1 let's
2 go   
3 to   
4 the  
5 q&a       
6 about
7 b&b
8 it's
9 great    

有没有办法向unnest_tokens() 发送一个选项来执行此操作,或者发送它当前使用的正则表达式并手动调整它以不包含与号?

【问题讨论】:

    标签: r tokenize tidytext unnest


    【解决方案1】:

    我们可以将token 用作regex

    library(tidytext)
    library(dplyr)
    d %>% 
       unnest_tokens(word, txt, token="regex", pattern = "[\\s!,.]")
    # A tibble: 9 x 1
    #  word 
    #  <chr>
    #1 let's
    #2 go   
    #3 to   
    #4 the  
    #5 q&a  
    #6 about
    #7 b&b  
    #8 it's 
    #9 great
    

    【讨论】:

    • 这行得通,但它也会留下标点符号(例如,如果我们在另一个句子中添加,它会带上句号)。 token="words" 的标点符号删除非常好。你认为我最好的选择是通过 token="regex" 发送模式沿着 = "[\\s,.]" 行吗?
    • @RayVelcoro 你能用那个新案例更新你的帖子,以便我可以测试它
    • @RayVelcoro 似乎可以工作unnest_tokens(word, txt, token="regex", pattern = "[ ,.]"),但它可能需要更多测试用例
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2018-03-11
    • 1970-01-01
    • 1970-01-01
    • 2017-09-21
    • 2018-09-11
    • 1970-01-01
    • 2013-06-19
    相关资源
    最近更新 更多