【问题标题】:Unable to recognize text content from GCS uri using @google-cloud/speech无法使用 @google-cloud/speech 识别来自 GCS uri 的文本内容
【发布时间】:2020-09-01 05:49:31
【问题描述】:

当我从本地缓冲区加载文件时它正在工作。但是当我使用 GCS URI 加载相同的文件时,响应为空。

    const fileName = './audio.wav';

    // Reads a local audio file and converts it to base64
    const file = fs.readFileSync(fileName);
    const audioBytes = file.toString('base64');

    const audio = {
        uri: 'gs://bucket-name/path-to-audio/audio.wav'
        // content: audioBytes
    };
    const config = {
        audioChannelCount: 1,
        encoding: 'LINEAR16',
        sampleRateHertz: 16000,
        languageCode: 'ta-IN',
    };
    const request = {
        audio: audio,
        config: config,
    };

    // Detects speech in the audio file
    const [operation] = await client.longRunningRecognize(request);
    console.info('OPERATION STATUS', operation.name);

当我尝试使用 GCS URI 加载它时,我得到 null 作为响应。然而,当我尝试发送与缓冲区相同的文件时,我得到了正确的响应。

# from GCS
TRANSLATION STATUS true
OPERATION COMPLETE STATUS  3489419937829075659 null undefined

# from local file
TRANSLATION STATUS true
OPERATION COMPLETE STATUS  390578141483807025 வணக்கம் வணக்கம் வணக்கம் SpeechRecognitionResult {
  alternatives: [
    SpeechRecognitionAlternative {
      words: [],
      transcript: 'வணக்கம் வணக்கம் வணக்கம்',
      confidence: 0.8997038006782532
    }
  ]
}

当我做一个控制台操作名 *3489419937829075659 *

STATUS DATA Operation {
  _events: [Object: null prototype] {
    newListener: [Function],
    removeListener: [Function]
  },
  _eventsCount: 2,
  _maxListeners: undefined,
  completeListeners: 0,
  hasActiveListeners: false,
  latestResponse: {
    name: '3489419937829075659',
    metadata: {
      type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeMetadata',
      value: <Buffer 08 64 12 0c 08 d7 9e b7 fa 05 10 90 d6 a3 a5 02 1a 0c 08 dc 9e b7 fa 05 10 e0 b7 f0 c8 02 22 40 67 73 3a 2f 2f 73 74 61 67 69 6e 67 2e 63 65 72 74 69 ... 46 more bytes>
    },
    done: true,
    response: {
      type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeResponse',
      value: <Buffer >
    },
    result: 'response'
  },
  name: '3489419937829075659',
  done: true,
  error: undefined,
  longrunningDescriptor: LongRunningDescriptor {
    operationsClient: OperationsClient {
      auth: [GoogleAuth],
      innerApiCalls: [Object],
      descriptor: [Object]
    },
    responseDecoder: [Function: bound decode_setup],
    metadataDecoder: [Function: bound decode_setup]
  },
  result: LongRunningRecognizeResponse { results: [] },
  metadata: LongRunningRecognizeMetadata {
    progressPercent: 100,
    startTime: Timestamp { seconds: [Long], nanos: 615050000 },
    lastUpdateTime: Timestamp { seconds: [Long], nanos: 689708000 }
  },
  backoffSettings: {
    initialRetryDelayMillis: 100,
    retryDelayMultiplier: 1.3,
    maxRetryDelayMillis: 60000,
    initialRpcTimeoutMillis: null,
    rpcTimeoutMultiplier: null,
    maxRpcTimeoutMillis: null,
    totalTimeoutMillis: null
  },
  response: {
    type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeResponse',
    value: <Buffer >
  },
  _callOptions: undefined,
  [Symbol(kCapture)]: false
}
STATUS DATA undefined

对于当我将整个对象进行控制台操作时的操作,我得到了这个,

STATUS DATA Operation {
  _events: [Object: null prototype] {
    newListener: [Function],
    removeListener: [Function]
  },
  _eventsCount: 2,
  _maxListeners: undefined,
  completeListeners: 0,
  hasActiveListeners: false,
  latestResponse: {
    name: '390578141483807025',
    metadata: {
      type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeMetadata',
      value: <Buffer 08 64 12 0c 08 f4 a3 b7 fa 05 10 88 e0 d4 ab 02 1a 0c 08 f7 a3 b7 fa 05 10 88 b0 ef d4 01>
    },
    done: true,
    response: {
      type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeResponse',
      value: <Buffer 12 4a 0a 48 0a 41 e0 ae b5 e0 ae a3 e0 ae 95 e0 af 8d e0 ae 95 e0 ae ae e0 af 8d 20 e0 ae b5 e0 ae a3 e0 ae 95 e0 af 8d e0 ae 95 e0 ae ae e0 af 8d 20 ... 26 more bytes>
    },
    result: 'response'
  },
  name: '390578141483807025',
  done: true,
  error: undefined,
  longrunningDescriptor: LongRunningDescriptor {
    operationsClient: OperationsClient {
      auth: [GoogleAuth],
      innerApiCalls: [Object],
      descriptor: [Object]
    },
    responseDecoder: [Function: bound decode_setup],
    metadataDecoder: [Function: bound decode_setup]
  },
  result: LongRunningRecognizeResponse {
    results: [ [SpeechRecognitionResult] ]
  },
  metadata: LongRunningRecognizeMetadata {
    progressPercent: 100,
    startTime: Timestamp { seconds: [Long], nanos: 628437000 },
    lastUpdateTime: Timestamp { seconds: [Long], nanos: 446421000 }
  },
  backoffSettings: {
    initialRetryDelayMillis: 100,
    retryDelayMultiplier: 1.3,
    maxRetryDelayMillis: 60000,
    initialRpcTimeoutMillis: null,
    rpcTimeoutMultiplier: null,
    maxRpcTimeoutMillis: null,
    totalTimeoutMillis: null
  },
  response: {
    type_url: 'type.googleapis.com/google.cloud.speech.v1.LongRunningRecognizeResponse',
    value: <Buffer 12 4a 0a 48 0a 41 e0 ae b5 e0 ae a3 e0 ae 95 e0 af 8d e0 ae 95 e0 ae ae e0 af 8d 20 e0 ae b5 e0 ae a3 e0 ae 95 e0 af 8d e0 ae 95 e0 ae ae e0 af 8d 20 ... 26 more bytes>
  },
  _callOptions: undefined,
  [Symbol(kCapture)]: false
}
STATUS DATA SpeechRecognitionResult {
  alternatives: [
    SpeechRecognitionAlternative {
      words: [],
      transcript: 'வணக்கம் வணக்கம் வணக்கம்',
      confidence: 0.8997038006782532
    }
  ]
}

【问题讨论】:

  • 您能发布您的完整代码吗?因为您当前的记录在这两种情况下都只记录操作名称console.info('OPERATION STATUS', operation.name);,而且据我所知,它确实在工作并记录数字3489419937829075659 和390578141483807025
  • 以上是执行操作的代码。

标签: node.js google-cloud-platform google-cloud-storage google-cloud-speech


【解决方案1】:

此代码适用于本地和 GCS:

    async function main() {
      // Imports the Google Cloud client library
      const speech = require('@google-cloud/speech');
      const fs = require('fs');
    
      // Creates a client
      const client = new speech.SpeechClient();
    
      // The name of the audio file to transcribe
      const fileName = './audio.wav';
    
      // Reads a local audio file and converts it to base64
      const file = fs.readFileSync(fileName);
      const audioBytes = file.toString('base64');
    
      // The audio file's encoding, sample rate in hertz, and BCP-47 language code
      const audio = {
          uri: "gs://BUCKET_NAME/audio.wav"
        //content: audioBytes,
      };
      // these config could be different from one audio type to another
      const config = {
        audioChannelCount: 1,
        encoding: 'LINEAR16',
        sampleRateHertz: 8000,
        languageCode: 'en-US',
      };
      const request = {
        audio: audio,
        config: config,
      };
    
      // Detects speech in the audio file
      const [operation] = await client.longRunningRecognize(request);
      const [response] = await operation.promise();
      const transcription = response.results
        .map(result => result.alternatives[0].transcript)
        .join('\n');
      console.log(`Transcription: ${transcription}`);
    }
    
    main().catch(console.error);

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多