Capabilities
Audio formats and endpoints
Which audio encodings, sample rates, and file sizes the STT API accepts on the synchronous and WebSocket endpoints.
Supported audio formats
| Format | Sync API | WebSocket Streaming |
|---|---|---|
LINEAR16 (PCM) | ||
MP3 | ||
OGG_OPUS | ||
FLAC | ||
AUTO_DETECT |
Recommended defaults: 16,000 Hz sample rate, 16-bit depth, mono. For container formats (MP3, FLAC, OGG_OPUS, WAV), sampleRateHertz is optional — the API auto-detects it from the file header.
Sync transcription accepts audio files up to ~16 MB. The actual duration depends on the encoding — for example, ~18 minutes of MP3 or ~8 minutes of 16 kHz 16-bit WAV. For larger files, split them into chunks or use the WebSocket streaming endpoint.
STT performs best with 16 kHz audio. Lower sample rates (such as 8 kHz telephony audio) contain fewer data points for the model to interpret, which reduces transcription accuracy. Upsampling low-sample-rate audio does not improve quality — it only interpolates between existing samples without adding new information.
Endpoints
| Endpoint | Method | Description |
|---|---|---|
/stt/v1/transcribe | POST | Send complete audio, receive full transcript |
/stt/v1/transcribe:streamBidirectional | WebSocket | Stream audio in real time, receive transcription chunks as they become available |