Skip to content

Best practices

Important

This document applies only to the China (Beijing) region. To use the model, you must use an API key for the China (Beijing) region.

Preprocess video files to improve file transcription efficiency (for audio file recognition scenarios)

Paraformer speech recognition API is compatible with video files, but they are typically large and time-consuming to transfer. Pre-process video files by extracting the audio track needed for speech recognition and compressing it to significantly reduce file size and improve transcription throughput. Use ffmpeg for pre-processing.

Prerequisites

Install ffmpeg from ffmpeg.org.

Pre-process video files

Use ffmpeg to extract the first audio track, downsample to 16kHz, and compress with opus encoding. Shell

HELPCODEESCAPE-shell
ffmpeg -i input-video-file -ac 1 -ar 16000 -acodec libopus output-audio-file.opus

The output audio file will be significantly smaller than the input video. Submit the audio file (via URL) to the file transcription API for speech recognition results.

Mirror of Alibaba Cloud Model Studio docs for reference and RAG. Not affiliated with Alibaba Cloud.