Get Started
Intro to Realtime TTS
Generate natural, expressive speech in real time.
Inworld's Realtime TTS models offer ultra-realistic, context-aware speech synthesis, zero data retention, and precise voice cloning capabilities, enabling developers to build natural and engaging experiences with human-like speech quality at an accessible price point.
Our models can be accessed via API (streaming and non-streaming) or the TTS Playground.
Developer quickstart
Learn how to make your first API call with a guided tutorial.
TTS Playground
Try different TTS models and voice cloning in TTS Playground.
Code Examples
Browse ready-to-use GitHub samples for common use cases.
Using AI to code? Paste https://docs.inworld.ai/llms.txt into your assistant so it knows every page on this site. Want live search? Add the MCP server.
Prefer the terminal? npm install -g @inworld/cli — synthesize speech, clone and design voices, and generate audiobooks with the Inworld CLI. AI agents can use it too.
Models
Realtime TTS-2
Our flagship, top-ranked model — the best choice for production
- Best quality and steerability, with natural language steering for more contextually aware speech
- Support for 200+ languages and locales
- Ultra-low latency (100 ms TTFB*) at high concurrency
- High quality instant voice cloning
- Enhanced timestamps with phonetic details and visemes
* P90 time to first audio byte, measured server-side — excludes network latency.
See the Models page for model IDs and full details.
Features
| Feature | Realtime TTS-2 | Realtime TTS-2 Flash |
|---|---|---|
| Best for | Production workloads that need the best quality and steerability | Latency-critical, high-volume, and cost-sensitive workloads |
| Quality | Top-ranked flagship — best quality and steerability | High quality at the lowest latency and cost |
| Latency (TTFB) * | 100 ms | 20 ms |
| Instant voice cloning | ||
| Professional voice cloning Beta | ||
| Inline pronunciation | ||
| Multilingual | 200+ languages | 200+ languages |
| Steering | ||
| Pause controls | ||
| Timestamp alignment | ||
| Zero data retention |
* P90 time to first audio byte, measured server-side — excludes network latency.