Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Best practices

STT best practices

Get the best transcripts from the STT API: audio choices, turn-taking, speech events, Voice Profile, language hints, and custom vocabulary.

  • Audio — Use MP3/OGG_OPUS for file uploads to reduce size; use LINEAR16 for streaming (required) and when you need highest quality.
  • Streaming — With manual turn-taking, send endTurn at turn boundaries while continuing the session. When done, send closeStream once to finalize pending audio and receive usage; do not precede it with another endTurn. See the WebSocket integration guide for the complete sequence.
  • Speech events — Listen for speechStarted and speechStopped events in the streaming response to detect when a speaker begins and stops talking. Use these to build custom turn-taking logic or visualize voice activity.
  • Voice Profile — Set voiceProfileConfig.enableVoiceProfile to true and optionally adjust topN (default: 10) to control how many labels per category are returned.
  • Language hint — If you know the audio's language, set language (e.g. en, ja) for cleaner output; the hint also constrains the output script (see Language Support). Leave it empty to auto-detect or when speakers switch languages.
  • Custom vocabulary — Pass domain-specific terms (names, jargon, acronyms) in prompts to bias recognition toward them. It is a soft bias rather than a hard keyword lock, so test it on the cases where the baseline actually misses the term.
  • Test with sample audio and your target language before production.