Developer quickstart
Learn how to make your first API call with a guided tutorial.
TTS Playground
Try different TTS models and voice cloning in TTS Playground.
Code Examples
Browse ready-to-use GitHub samples for common use cases.
Models
Realtime TTS-2
Our flagship, top-ranked model — the best choice for production
- Best quality and steerability, with natural language steering for more contextually aware speech
- Support for 200+ languages and locales
- Ultra-low latency (100 ms TTFB*) at high concurrency
- High quality instant voice cloning
- Enhanced timestamps with phonetic details and visemes
Realtime TTS-2 Flash
Our fastest, most cost-efficient model — built for latency-critical, high-volume workloads
- Our lowest latency — 20 ms TTFB*, 5× faster than
inworld-tts-2 - Lowest cost per character
- Same 200+ languages and locales as
inworld-tts-2 - High quality instant voice cloning
Features
* P90 time to first audio byte, measured server-side — excludes network latency.