【发布时间】:2017-02-03 22:00:20
【问题描述】:
在处理具有需要文件路径作为参数的方法的 Python 库时,我经常遇到问题。当我在内存中有一些我想与库函数一起使用的数据时,这是一个问题。在这些情况下,我最终会做的是:
- 写入包含数据的临时文件。
- 将临时文件路径传递给库函数。
- 函数返回后删除文件。
这工作得很好,但是,对于时间敏感的应用程序,涉及到临时文件的写入和读取的文件 IO 是一个交易破坏者。
有没有人可以解决这个问题?我认为这里没有一种适合所有解决方案的解决方案,但我不想做任何假设。但是,让我描述一下我当前的用例,希望有人能够具体帮助我。
我正在使用speech_recognition 库将大量音频文件转换为文本。我有二进制形式的音频文件数据。这是我的代码:
from os import path, remove
from scipy.io.wavfile import write
import speech_recognition as sr
audio_list = ... # get the audio
text_list = []
for item in audio_list:
temp_name = 'temp.wav'
# create temporary file, writing it as a wave for speech_recognition to read
write(temp_name, rate, item)
audio_file = path.join(path.dirname(path.realpath('__file__')), temp_name)
recognizer = sr.Recognizer()
# this is where I need to have the path to the file
with sr.AudioFile(audio_file) as source:
audio = recognizer.record(source)
text = recognizer.recognize_sphinx(audio)
text_list.append(text)
remove(temp_name)
speech_recognition 库使用PocketSphinx 作为后端。 PocketSphinx 有自己的 Python API,但我也没有运气。
谁能帮我减少这个文件的IO?
【问题讨论】:
标签: python python-3.x speech-recognition speech-to-text cmusphinx