【问题标题】:How to avoid punctuationduring tokenization using Stanford NLP如何在使用斯坦福 NLP 进行标记化时避免标点符号
【发布时间】:2015-03-15 13:51:22
【问题描述】:

我正在使用斯坦福核心 NLP。我已经尝试了以下示例。这个例子可以标记文本中的单词。但是它也提取标点符号,例如逗号,句号等。我想知道如何设置不允许提取标点符号的属性,或者是否有其他方法可以做到这一点。这是代码示例。我知道使用 Python 很容易,但不知道如何在 Java 中做到这一点。请提出建议。

    props = new Properties();
    props.setProperty("annotators", "tokenize, ssplit");
    pipeline = new StanfordCoreNLP(props);
    String text = "this is simple text written in English,Spanish etc."

// create an empty Annotation just with the given text
    Annotation document = new Annotation(text);

   pipeline.annotate(document);

   List<CoreMap> sentences = document.get(SentencesAnnotation.class);

   for(CoreMap sentence: sentences) {
     for (CoreLabel token: sentence.get(TokensAnnotation.class)) {
    // this is the text of the token
    String word = token.get(TextAnnotation.class);
      }
   }

【问题讨论】:

    标签: tokenize stanford-nlp punctuation


    【解决方案1】:

    我们没有任何标记器选项可以跳过这些,但这应该不难。标点字符串是封闭类。

    您可以使用正则表达式匹配标点符号。 (使用\p{Punct};参见例如Punctuation Regex in Java)。然后只需删除其文本内容与此类正则表达式匹配的标记。

    【讨论】:

    • 谢谢,停用词呢,如何丢弃停用词,它们也不能作为选项使用吗?
    • 它们不能作为标准功能使用,因为停用词的定义不是很具体/它是特定于任务的。您可以建立自己的停用词列表或在 Internet 上找到一个并手动过滤停用词令牌,方法与使用标点符号相同。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-10-12
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多