Best practices
TTS best practices
Short answers to common questions about making Inworld TTS expressive, fast, and ready for production.
Teams building with Inworld TTS ask these three questions most often. Start with the answers below, then follow the links for more detail.
How do I make it as expressive as possible?
- Start from the right voice. The models are designed to closely follow the voice sample. A calm, flat sample produces calm, flat speech no matter what the text says. Pick or clone a voice whose sample has the energy and emotional range you want.
- Write for the ear. Punctuation, capitalization for emphasis, and filler words shape delivery more than any parameter. Speech text is not display text.
- Use non-verbal sounds. Inline tags like
[laugh],[sigh], and[breathe]make speech sound more human. They work on bothinworld-tts-2andinworld-tts-2-flash. - Steer delivery on
inworld-tts-2. Instruction tags like[say excitedly with a fast pace]guide emotion, pace, and vocal style. Use them where delivery should change. You don't need them on every sentence. - Give the model context. In
synthesisContext, pass the text of earlier requests from the same conversation. This helps short replies like "Yeah." use the right intonation. Keep it to the last few turns. Context text is not billed.
Generating Naturally Sounding Speech
The full guide: voice selection, writing for speech, non-verbals, steering, context, normalization, and LLM prompting.
How do I make it as fast as possible?
- Pick the right model for your latency budget.
inworld-tts-2-flashdelivers 20 ms time to first audio byte (P90, measured server-side) — 5× faster thaninworld-tts-2at 100 ms.Prefer
inworld-tts-2for the best quality and realism. Its 100 ms TTFB is already fast enough for realtime applications. Chooseinworld-tts-2-flashonly when speed matters most and you don't need steering or professional voice cloning. - Stream everything. Use WebSocket streaming (lowest latency) or HTTP streaming and start playback on the first chunk. If an LLM generates your text, stream its output into TTS sentence by sentence instead of waiting for the full response.
- Turn off text normalization. Server-side normalization adds latency to every request. For latency-sensitive applications, set
applyTextNormalizationtoOFFand have your LLM or application write numbers, dates, and symbols in spoken form instead. - Run your backend nearby. Time to first audio includes the network round trip between your servers and ours. Run your backend as close to Inworld as your infrastructure allows. Inworld has a default US deployment and regional deployments in the EU and India. If you need another region, contact us.
- Reuse connections. Keep connections alive between requests to skip repeated TCP/TLS handshakes, and don't wait for acknowledgments the protocol doesn't require.
Latency
The full guide: streaming setups, connection reuse, chunking, and WebSocket-specific techniques.
How do I get ready for my production launch?
- Choose the right plan. Concurrency limits range from 5 concurrent generations on On-Demand to 500 on Growth. Enterprise has custom limits. Compare tiers on the pricing page and make sure your plan covers your expected peak before launch.
- Estimate the concurrency you need. Speech generates much faster than it plays, so one generation slot typically serves several conversations at once. We usually see at least 4× more conversations than slots. Use your traffic patterns to estimate how many slots you need.
- Retry limit errors. Requests over your concurrency limit are best-effort and may be rejected with a
429(HTTP) orcode: 8error (WebSocket). Use retries with backoff to handle these errors. - Choose an API that fits the input. Non-streaming HTTP accepts up to 2,000 characters per request, streaming HTTP up to 4,000 characters per request, and each WebSocket
send_textmessage up to 2,000 characters. For articles, chapters, and other long content, start with the Async API, which accepts up to 100,000 characters per job (10,000 on On-Demand). See Long Text Input for details and client-side chunking options.