Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Best practices

TTS best practices

Short answers to common questions about making Inworld TTS expressive, fast, and ready for production.

Teams building with Inworld TTS ask these three questions most often. Start with the answers below, then follow the links for more detail.

How do I make it as expressive as possible?

  1. Start from the right voice. The models are designed to closely follow the voice sample. A calm, flat sample produces calm, flat speech no matter what the text says. Pick or clone a voice whose sample has the energy and emotional range you want.
  2. Write for the ear. Punctuation, capitalization for emphasis, and filler words shape delivery more than any parameter. Speech text is not display text.
  3. Use non-verbal sounds. Inline tags like [laugh], [sigh], and [breathe] make speech sound more human. They work on both inworld-tts-2 and inworld-tts-2-flash.
  4. Steer delivery on inworld-tts-2. Instruction tags like [say excitedly with a fast pace] guide emotion, pace, and vocal style. Use them where delivery should change. You don't need them on every sentence.
  5. Give the model context. In synthesisContext, pass the text of earlier requests from the same conversation. This helps short replies like "Yeah." use the right intonation. Keep it to the last few turns. Context text is not billed.

Generating Naturally Sounding Speech

The full guide: voice selection, writing for speech, non-verbals, steering, context, normalization, and LLM prompting.

How do I make it as fast as possible?

  1. Pick the right model for your latency budget. inworld-tts-2-flash delivers 20 ms time to first audio byte (P90, measured server-side) — 5× faster than inworld-tts-2 at 100 ms.

    Prefer inworld-tts-2 for the best quality and realism. Its 100 ms TTFB is already fast enough for realtime applications. Choose inworld-tts-2-flash only when speed matters most and you don't need steering or professional voice cloning.

  2. Stream everything. Use WebSocket streaming (lowest latency) or HTTP streaming and start playback on the first chunk. If an LLM generates your text, stream its output into TTS sentence by sentence instead of waiting for the full response.
  3. Turn off text normalization. Server-side normalization adds latency to every request. For latency-sensitive applications, set applyTextNormalization to OFF and have your LLM or application write numbers, dates, and symbols in spoken form instead.
  4. Run your backend nearby. Time to first audio includes the network round trip between your servers and ours. Run your backend as close to Inworld as your infrastructure allows. Inworld has a default US deployment and regional deployments in the EU and India. If you need another region, contact us.
  5. Reuse connections. Keep connections alive between requests to skip repeated TCP/TLS handshakes, and don't wait for acknowledgments the protocol doesn't require.

Latency

The full guide: streaming setups, connection reuse, chunking, and WebSocket-specific techniques.

How do I get ready for my production launch?

  1. Choose the right plan. Concurrency limits range from 5 concurrent generations on On-Demand to 500 on Growth. Enterprise has custom limits. Compare tiers on the pricing page and make sure your plan covers your expected peak before launch.
  2. Estimate the concurrency you need. Speech generates much faster than it plays, so one generation slot typically serves several conversations at once. We usually see at least 4× more conversations than slots. Use your traffic patterns to estimate how many slots you need.
  3. Retry limit errors. Requests over your concurrency limit are best-effort and may be rejected with a 429 (HTTP) or code: 8 error (WebSocket). Use retries with backoff to handle these errors.
  4. Choose an API that fits the input. Non-streaming HTTP accepts up to 2,000 characters per request, streaming HTTP up to 4,000 characters per request, and each WebSocket send_text message up to 2,000 characters. For articles, chapters, and other long content, start with the Async API, which accepts up to 100,000 characters per job (10,000 on On-Demand). See Long Text Input for details and client-side chunking options.