【问题标题】:How to keep special symbols like "(" "," and "#" in tokens in R?如何在 R 的标记中保留特殊符号,如“(”、“、”和“#”?
【发布时间】:2018-03-11 00:25:12
【问题描述】:

我正在处理一个文本文件,其中包含来自招聘广告的“c#”、“c++”和“.net”等词。当我将其转换为标记时,“#”、“++”和点被删除。如何将它们保留在生成的令牌中?这是我的代码:

unnest_tokens(word,REQUIREMENTS, token = "words",to_lower=TRUE)

【问题讨论】:

    标签: r data-mining tokenize


    【解决方案1】:

    问题在于参数token = "words",它在非单词字符上进行拆分(可能使用正则表达式\\W+)。此函数会丢弃分隔符,因此为了保留这些字符,您必须使用"words" 以外的其他参数。您可能希望使用 token = "regex" 定义自己的拆分正则表达式,如下所示:

    unnest_tokens(word,
                  REQUIREMENTS,
                  token = "regex",
                  to_lower = TRUE,
                  pattern = "\\s+") # split on whitespace rather than non-word elements
    

    这样,you can define whatever regex you need 可以自定义文本的标记方式。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2011-05-08
      • 1970-01-01
      • 2021-05-14
      • 1970-01-01
      • 2020-08-04
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多