Realtime TTS-2 is live. Built for realtime conversation that feels human. Learn more

Get Started

Intro to Realtime TTS

Generate natural, expressive speech in real time.

Inworld's Realtime TTS models offer ultra-realistic, context-aware speech synthesis, zero data retention, and precise voice cloning capabilities, enabling developers to build natural and engaging experiences with human-like speech quality at an accessible price point.

Our models can be accessed via API (streaming and non-streaming) or the TTS Playground.

Using AI to code? Paste https://docs.inworld.ai/llms.txt into your assistant so it knows every page on this site. Want live search? Add the MCP server.

Prefer the terminal? npm install -g @inworld/cli — synthesize speech, clone and design voices, and generate audiobooks with the Inworld CLI. AI agents can use it too.

Models

Realtime TTS-2

Our flagship, top-ranked model — the best choice for production

  • Best quality and steerability, with natural language steering for more contextually aware speech
  • Support for 200+ languages and locales
  • Ultra-low latency (100 ms TTFB*) at high concurrency
  • High quality instant voice cloning
  • Enhanced timestamps with phonetic details and visemes

Realtime TTS-2 Flash

Our fastest, most cost-efficient model — built for latency-critical, high-volume workloads

  • Our lowest latency — 20 ms TTFB*, 5× faster than inworld-tts-2
  • Lowest cost per character
  • Same 200+ languages and locales as inworld-tts-2
  • High quality instant voice cloning

* P90 time to first audio byte, measured server-side — excludes network latency.

See the Models page for model IDs and full details.

Features

FeatureRealtime TTS-2Realtime TTS-2 Flash
Best for                Production workloads that need the best quality and steerabilityLatency-critical, high-volume, and cost-sensitive workloads
Quality                Top-ranked flagship — best quality and steerabilityHigh quality at the lowest latency and cost
Latency (TTFB) *                100 ms20 ms
Instant voice cloning                
Professional voice cloning Beta                
Inline pronunciation                
Multilingual                200+ languages200+ languages
Steering                
Pause controls                
Timestamp alignment                
Zero data retention                

* P90 time to first audio byte, measured server-side — excludes network latency.