【问题标题】:Finding start and end point of sentence in a paragraph StanfordCoreNLP在段落中查找句子的开始和结束点 StanfordCoreNLP
【发布时间】:2016-02-10 00:06:36
【问题描述】:

我想知道如何使用 StanfordCoreNLP 在段落中找到句子的开始和结束位置。现在我正在使用 DocumentPreprocessor 将段落拆分为句子。是否可以得到句子在原文中实际所在位置的起止索引?

我正在使用此处提出的另一个问题的代码。

String paragraph = "My 1st sentence. “Does it work for questions?” My third sentence.";
Reader reader = new StringReader(paragraph);
DocumentPreprocessor dp = new DocumentPreprocessor(reader);
List<String> sentenceList = new ArrayList<String>();

for (List<HasWord> sentence : dp) {
   String sentenceString = Sentence.listToString(sentence);
   sentenceList.add(sentenceString.toString());
}

for (String sentence : sentenceList) {
   System.out.println(sentence);
}

取自:How can I split a text into sentences using the Stanford parser?

谢谢

【问题讨论】:

    标签: java indexing split stanford-nlp sentence


    【解决方案1】:

    快速而肮脏的方法是:

    import edu.stanford.nlp.simple.*;
    
    Document doc = new Document("My 1st sentence. “Does it work for questions?” My third sentence.");
    for (Sentence sentence : doc.sentences()) {
      System.out.println(sentence.characterOffsetBegin(0) + " -- " + sentence.characterOffsetEnd(sentence.length() - 1));
    }
    

    否则,您可以从 CoreLabel 中提取 CharacterOffsetBeginAnnotation 和 CharacterOffsetEndAnnotation,并使用它来查找令牌在原始文本中的偏移量。

    【讨论】:

      【解决方案2】:

      有关获取 CharacterOffsetEndAnnotation 的示例,请参阅 https://www.programcreek.com/java-api-examples/?api=edu.stanford.nlp.ling.CoreLabel

      【讨论】:

      • 链接答案没问题,但通常您需要描述链接的内容。这是因为您链接的网站可能会关闭,这意味着您的回答将毫无用处。
      猜你喜欢
      • 2012-10-18
      • 2023-04-02
      • 2016-07-22
      • 1970-01-01
      • 2022-01-10
      • 1970-01-01
      • 2019-07-06
      • 2023-03-30
      相关资源
      最近更新 更多