> ## Documentation Index
> Fetch the complete documentation index at: https://docs.inworld.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio formats and endpoints

> Which audio encodings, sample rates, and file sizes the STT API accepts on the synchronous and WebSocket endpoints.

## Supported audio formats

| **Format** | **Sync API** | **WebSocket Streaming** |
| :--- | :--- | :--- |
| `LINEAR16` (PCM) | <Icon icon="check" size={18} /> | <Icon icon="check" size={18} /> |
| `MP3` | <Icon icon="check" size={18} /> | <Icon icon="xmark" size={18} /> |
| `OGG_OPUS` | <Icon icon="check" size={18} /> | <Icon icon="xmark" size={18} /> |
| `FLAC` | <Icon icon="check" size={18} /> | <Icon icon="xmark" size={18} /> |
| `AUTO_DETECT` | <Icon icon="check" size={18} /> | <Icon icon="xmark" size={18} /> |

Recommended defaults: 16,000 Hz sample rate, 16-bit depth, mono. For container formats (MP3, FLAC, OGG_OPUS, WAV), `sampleRateHertz` is optional — the API auto-detects it from the file header.

<Note>
Sync transcription accepts audio files up to **~16 MB**. The actual duration depends on the encoding — for example, ~18 minutes of MP3 or ~8 minutes of 16 kHz 16-bit WAV. For larger files, split them into chunks or use the [WebSocket streaming endpoint](https://docs.inworld.ai/api-reference/sttAPI/speechtotext/transcribe-stream-websocket.md).
</Note>

<Note>
STT performs best with 16 kHz audio. Lower sample rates (such as 8 kHz telephony audio) contain fewer data points for the model to interpret, which reduces transcription accuracy. Upsampling low-sample-rate audio does not improve quality — it only interpolates between existing samples without adding new information.
</Note>

## Endpoints

| **Endpoint** | **Method** | **Description** |
| :--- | :--- | :--- |
| [`/stt/v1/transcribe`](https://docs.inworld.ai/api-reference/sttAPI/speechtotext/transcribe.md) | POST | Send complete audio, receive full transcript |
| [`/stt/v1/transcribe:streamBidirectional`](https://docs.inworld.ai/api-reference/sttAPI/speechtotext/transcribe-stream-websocket.md) | WebSocket | Stream audio in real time, receive transcription chunks as they become available |
