【问题标题】:Google Cloud Platform speech to text - customize transcribed textGoogle Cloud Platform 语音转文本 - 自定义转录文本
【发布时间】:2021-07-01 20:40:00
【问题描述】:

我正在使用 Google Cloud Platform (GCP) model adaptation feature of speech to text 来识别行业独有的话语,例如当用户说出 JSON 时,它应该被转录为 JSON 而不是“Jason”。我通过使用短语集和相关的提升值来实现这一点。

本示例中的文本转录为 Json。我希望将其转录为 JSON(全部大写)

我已通读 GCP 文档,但没有找到与我的问题相关的文档。我也试过 Azure,there's an option to upload a pronunciation file。我正在 GCP 中寻找类似的解决方案。

【问题讨论】:

    标签: google-cloud-platform speech-to-text google-cloud-speech


    【解决方案1】:

    我自己试过了,得到了同样的结果。即使maxAlternatives 设置为 20。

    目前没有像发音文件这样的选项,所以我创建了一个Feature Request 来请求它的实现。
    记得给它加星标,以便在每次更新时收到电子邮件通知。并且,如果可以,请添加您的业务案例和/或影响以提供完整的图片。

    目前,解决方法是在您的代码上实现“捕手”。 在 Python 中,您可以使用 replace() 或 upper()。
    类似的东西:

    for result in response.results:
        print("Transcript: {}".format(result.alternatives[0].transcript.replace('Json', 'JSON')))
    

    如果您需要捕获更多单词,请使用 if 条件遍历列表:

    result='I need a Json file'
    lower_words = ['Json', 'csv']
    upper_words = ['JSON', 'CSV']
    for result in response.results:
        for lower_word, upper_word in zip(lower_words, upper_words):
            if lower_word in result:
                print("Transcript: {}".format(result.alternatives[0].transcript.replace(lower_word, upper_word)))
    

    当然,这将在满足条件的每次迭代时打印,因此如果结果中可能包含多个单词,您可能希望存储中间结果并在嵌套循环之后打印。

    我希望您不要更改太多单词,否则这会大大降低您的应用程序速度。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-01-27
      • 1970-01-01
      • 1970-01-01
      • 2012-09-29
      相关资源
      最近更新 更多