【发布时间】:2015-03-15 13:51:22
【问题描述】:
我正在使用斯坦福核心 NLP。我已经尝试了以下示例。这个例子可以标记文本中的单词。但是它也提取标点符号,例如逗号,句号等。我想知道如何设置不允许提取标点符号的属性,或者是否有其他方法可以做到这一点。这是代码示例。我知道使用 Python 很容易,但不知道如何在 Java 中做到这一点。请提出建议。
props = new Properties();
props.setProperty("annotators", "tokenize, ssplit");
pipeline = new StanfordCoreNLP(props);
String text = "this is simple text written in English,Spanish etc."
// create an empty Annotation just with the given text
Annotation document = new Annotation(text);
pipeline.annotate(document);
List<CoreMap> sentences = document.get(SentencesAnnotation.class);
for(CoreMap sentence: sentences) {
for (CoreLabel token: sentence.get(TokensAnnotation.class)) {
// this is the text of the token
String word = token.get(TextAnnotation.class);
}
}
【问题讨论】:
标签: tokenize stanford-nlp punctuation