Realtime TTS-2 is live. Built for realtime conversation that feels human. Learn more

Get Started

TTS Models

Inworld provides a family of state-of-the-art TTS models, optimized for different use cases, quality levels, and performance requirements.

Realtime TTS-2

Our flagship, top-ranked model — the best choice for production

  • Best quality and steerability, with natural language steering for more contextually aware speech
  • Support for 200+ languages and locales
  • Ultra-low latency (100 ms TTFB*) at high concurrency
  • High quality instant voice cloning
  • Enhanced timestamps with phonetic details and visemes

Realtime TTS-2 Flash

Our fastest, most cost-efficient model — built for latency-critical, high-volume workloads

  • Our lowest latency — 20 ms TTFB*, 5× faster than inworld-tts-2
  • Lowest cost per character
  • Same 200+ languages and locales as inworld-tts-2
  • High quality instant voice cloning

* P90 time to first audio byte, measured server-side — excludes network latency.

Models overview

NameModel IDDescriptionSupported languages
Realtime TTS-2inworld-tts-2              Our most powerful model with natural language steering and the strongest multilingual capabilities — the best choice for production quality200+ languages and locales — see Languages
Realtime TTS-2 Flashinworld-tts-2-flash              Our fastest, most cost-efficient model — significantly lower latency than inworld-tts-2, ideal for real-time agents and high-volume workloads200+ languages and locales — see Languages

Using an earlier model? Previous-generation models (inworld-tts-1, inworld-tts-1-max, inworld-tts-1.5-max, inworld-tts-1.5-mini) are deprecated and not recommended for new projects. inworld-tts-1 and inworld-tts-1-max were discontinued on June 15, 2026 and requests to them are automatically routed to newer models. If you're still using any of these, we recommend migrating to the TTS-2 family for better quality, latency, and cost.