Skip to main content
Inworld’s Realtime TTS models offer ultra-realistic, context-aware speech synthesis, zero data retention, and precise voice cloning capabilities, enabling developers to build natural and engaging experiences with human-like speech quality at an accessible price point. Our models can be accessed via API (streaming and non-streaming) or the TTS Playground.

Developer quickstart

Learn how to make your first API call with a guided tutorial.

TTS Playground

Try different TTS models and voice cloning in TTS Playground.

Code Examples

Browse ready-to-use GitHub samples for common use cases.
Using AI to code? Paste https://docs.inworld.ai/llms.txt into your assistant so it knows every page on this site. Want live search? Add the MCP server.Prefer the terminal? npm install -g @inworld/cli — synthesize speech, clone and design voices, and generate audiobooks with the Inworld CLI. AI agents can use it too.

Models

Realtime TTS-2

Our flagship, top-ranked model — the best choice for production

  • Best quality and steerability, with natural language steering for more contextually aware speech
  • Support for 200+ languages and locales
  • Ultra-low latency (100 ms TTFB*) at high concurrency
  • High quality instant voice cloning
  • Enhanced timestamps with phonetic details and visemes

Realtime TTS-2 Flash

Our fastest, most cost-efficient model — built for latency-critical, high-volume workloads

  • Our lowest latency — 20 ms TTFB*, 5× faster than inworld-tts-2
  • Lowest cost per character
  • Same 200+ languages and locales as inworld-tts-2
  • High quality instant voice cloning
* P90 time to first audio byte, measured server-side — excludes network latency. See the Models page for model IDs and full details.

Features

* P90 time to first audio byte, measured server-side — excludes network latency.