Realtime TTS-2 Flash
Launched Realtime TTS-2 Flash (inworld-tts-2-flash), the fastest member of the TTS-2 family — see Models:- Our lowest latency: 20 ms time to first audio (server-side P90 TTFB, excluding network latency) — 5× faster than
inworld-tts-2at 100 ms, making it the best choice for latency-critical real-time agents. - Our lowest cost: The most cost-efficient model per character, ideal for high-volume workloads.
- Full TTS-2 language coverage: The same 200+ languages and locales as
inworld-tts-2, plus instant voice cloning and timestamp alignment.
Steering instructions and Professional Voice Cloning are supported on
inworld-tts-2 only — use it when you need directed, contextually aware delivery. Non-verbal tags like [laugh] work on both models.Steering instructions now persist
Steering oninworld-tts-2 follows one rule: a [tag] applies from where you write it until you change it. See the Steering guide.[reset]: New reserved tag that ends a styled passage and returns the voice to its own character for the rest of the text.[shouting] We need to leave now! [reset] Do you understand me?shouts only the first sentence.- Instructions survive pauses: A
<break/>no longer clears the active instruction. A pause is a pause and never changes delivery. - Request-level
instructionfield: Set one instruction for the whole request without putting tags in your text. Seeinstruction. Use either this field or inline tags, not both.
Realtime TTS-2
Launched Realtime TTS-2 (inworld-tts-2), our most powerful and expressive TTS model:- Natural Language Steering: Direct any voice with bracketed instructions like
[say excitedly],[whisper in a hushed style], or free-form directions like[speak as if barely holding back rage]. Covers articulation, intonation, volume, pitch, range, speed, vocal style, and non-verbals ([laugh],[sigh], etc.). See the Steering guide. - Stronger Multilingual Support: Production-quality synthesis across 15 languages, plus experimental support for 90+ additional languages. See Languages.
- Cross-Lingual Voice Synthesis: Reuse the same voice across multiple languages. For best results, specify the
languagefield. - Voice Localization: Localize your voice for the most consistent, native-sounding speech in a target language. See Voice Localization.
- Delivery Mode: New
deliveryModefield (STABLE,BALANCED,CREATIVE) controls the trade-off between consistency and emotional range. - Updated Voice Design: Released an updated version of Voice Design with improved generations. See Voice Design.
Inworld TTS 1.5
Launched Inworld TTS 1.5, our newest generation of realtime TTS models featuring:- Two New Models: Our flagship model
inworld-tts-1.5-maxis ideal for most use cases, with the best balance of quality and speed. For use cases where latency is the top priority, we also offerinworld-tts-1.5-mini. - Latency Improvements: Our new TTS-1.5 models achieve P90 latency for first audio chunk delivery under 250ms for our Max model and under 130ms for our Mini model, a 4x improvement compared to TTS-1.
- More Expressive and More Stable: TTS-1.5 is 30% more expressive than prior generations and demonstrates a 40% reduction in word error rates.
- Additional Languages: We’ve added support for additional languages, including Hindi, Arabic, and Hebrew, bringing total languages supported to 15.
Updates to Inworld TTS
Released an upgraded version of the Inworld TTS models with higher overall quality.- Speech Quality: Clearer, more natural speech with smoother pacing and more accurate pronunciation.
- Voice Similarity: Cloned voices sound closer to the originals, preserving each voice’s unique style.
- Non-English Languages: More consistent, reliable output across supported non-English languages.
- Custom Pronunciation: New support for inline IPA, giving you control over exact word pronunciations. See the Key Features for details.