【发布时间】:2022-01-25 21:43:42
【问题描述】:
是否有一种可能的方法可以从没有任何标点符号和/或全部小写的段落中提取句子/句子标记化?我们特别需要能够将段落拆分为句子,同时预期输入的段落不正确的最坏情况。
例子:
this is a sentence this is a sentence this is a sentence this is a sentence this is a sentence
进入
["this is a sentence", "this is a sentence", "this is a sentence", "this is a sentence", "this is a sentence"]
到目前为止,我们尝试过的句子分词器似乎依赖于标点符号和真正的大小写:
使用 nltk.sent_tokenize
"This is a sentence. This is a sentence. This is a sentence"
进入
['This is a sentence.', 'This is a sentence.', 'This is a sentence']
【问题讨论】:
标签: python nlp nltk spacy linguistics