Realtime TTS-2
Our flagship, top-ranked model — the best choice for production
- Best quality and steerability, with natural language steering for more contextually aware speech
- Support for 200+ languages and locales
- Ultra-low latency (100 ms TTFB*) at high concurrency
- High quality instant voice cloning
- Enhanced timestamps with phonetic details and visemes
Realtime TTS-2 Flash
Our fastest, most cost-efficient model — built for latency-critical, high-volume workloads
- Our lowest latency — 20 ms TTFB*, 5× faster than
inworld-tts-2 - Lowest cost per character
- Same 200+ languages and locales as
inworld-tts-2 - High quality instant voice cloning
Models overview
Using an earlier model? Previous-generation models (
inworld-tts-1, inworld-tts-1-max, inworld-tts-1.5-max, inworld-tts-1.5-mini) are deprecated and not recommended for new projects. inworld-tts-1 and inworld-tts-1-max were discontinued on June 15, 2026 and requests to them are automatically routed to newer models. If you’re still using any of these, we recommend migrating to the TTS-2 family for better quality, latency, and cost.