【发布时间】:2020-09-28 19:42:55
【问题描述】:
我正在寻找可以尝试从单词序列中检测句子边界和结构的库或系统。这些单词来自演示文稿的抄本,包括填充词 um 和 uh 以及单词重复。这是一个例子:
["hello", "everybody", "today", "we", "are", "going", "to", "um", "be", "discussing", "IO", "streams", "with", "IO", "streams", "we", ...]
Google Cloud 能够使用 enableAutomaticPunctuation 选项在其 Speech-to-Text API 的文本输出中添加标点符号,但是我没有找到类似这样的将文本作为输入的东西。
【问题讨论】:
标签: nlp