Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

APIs

Synthesize speech (WebSocket)

Persistent connection for lowest-latency audio streaming

You open a persistent WebSocket connection and send text messages. The server streams audio chunks back over the same connection — no per-request overhead, no repeated handshakes. This gives you the lowest possible latency.

Best for voice agents and interactive applications that send multiple synthesis requests in a session, where avoiding connection setup on every call makes a measurable difference.

If you only need a single request-response with chunked audio, the Streaming API is simpler to integrate. For tips on optimizing latency, see the latency best practices guide.

Message pipelining

Messages on a connection are processed in order, so you don't need to wait for the contextCreated acknowledgment before sending text. For the lowest latency, send each message as soon as it is ready — waiting for the acknowledgment adds a full network round trip before the first audio chunk arrives. If you omit contextId, messages are automatically routed to the connection's auto-created context.

Choose when speech generation starts

Choose a mode based on how text arrives in your application:

Your inputChooseWho decides when to generate speech?
You have the full text, or want to control batching yourselfManual — autoMode: false (default)Your app calls flushContext, or closeContext when finished. Buffering safeguards also apply.
Your app already splits text into complete sentences or phrasesClient segmented — autoMode: true, autoModeStrategy: "CLIENT_SEGMENTED"The server schedules the segments your app sends.
You receive partial text or individual tokens from an LLMSentence boundary — autoMode: true, autoModeStrategy: "SENTENCE_BOUNDARY" (Preview)The server detects sentence boundaries in the incoming text.

A context holds the settings and text for a speech-generation session within the WebSocket connection. The buffer holds text waiting to be synthesized. To flush is to request synthesis of the buffered text, even if it has no final punctuation. Complete any markup tags before flushing.

Three text paths: manual text waits for your app to flush; client-split sentences are scheduled by the server; raw LLM tokens are buffered until the server detects a boundary. Each path leads to speech generation.

Set autoMode and autoModeStrategy in the create message when creating a context. Leave the buffering safeguards at their defaults for normal use.

Manual flushing

Use autoMode: false when you have a complete reply or want to decide how to batch text. Send text with sendText, then call flushContext to request speech generation. If you have no more text for this context, use closeContext instead.

To synthesize buffered text and keep the context open for more text, send:

javascript
ws.send(JSON.stringify({contextId: "reply-1", flushContext: {}}));

Use flushContext between batches you want to control yourself. When finished, close the context to release any remaining text and finish receiving audio.

autoModeStrategy is ignored in this mode. See the runnable manual-flushing guide, which also covers interruptions and keeping LLM history aligned with what the user heard.

Client-segmented streaming

Use autoMode: true with autoModeStrategy: "CLIENT_SEGMENTED" when your app already splits text into sentences or phrases. Send each complete segment as it becomes available. The server controls flushing and can combine queued segments while earlier audio is streaming.

Auto-mode default: autoMode: true on its own selects CLIENT_SEGMENTED. Omitting autoModeStrategy, or setting it to AUTO_MODE_STRATEGY_UNSPECIFIED, has the same effect. To send unsplit LLM tokens, explicitly select SENTENCE_BOUNDARY.

See the runnable client-side sentence-splitting guide.

Server sentence-boundary streaming

Use autoMode: true with autoModeStrategy: "SENTENCE_BOUNDARY" to send text as your LLM produces it. The server buffers incoming fragments and detects sentence boundaries and complete break tags.

Preview: SENTENCE_BOUNDARY supports inworld-tts-2 and inworld-tts-2-flash where those models are available. Selecting it with auto mode enabled on an unsupported model returns INVALID_ARGUMENT.

The tail is the remaining buffered text that has not reached a confirmed boundary. When your text source finishes, close the context to release this remaining text. If you need to release the tail and continue using the same context, you can call flushContext instead. Auto mode does not use the buffer delay timer.

See the runnable LLM token-streaming guide.

Stream incremental text

This example uses an already connected, authenticated WebSocket. Each ws.send sends one message. In your application, send each fragment as your LLM produces it.

javascript
ws.send(JSON.stringify({
  contextId: "reply-1",
  create: {
    voiceId: "Ashley",
    modelId: "inworld-tts-2",
    autoMode: true,
    autoModeStrategy: "SENTENCE_BOUNDARY"
  }
}));

for (const text of ["Hello from ", "sentence boundary streaming.", " Next we will finish this thought"]) {
  ws.send(JSON.stringify({contextId: "reply-1", sendText: {text}}));
}

// The LLM has finished: synthesize remaining text and close this context.
ws.send(JSON.stringify({contextId: "reply-1", closeContext: {}}));
// Keep receiving audio until the server sends contextClosed.

Here is how the buffer changes as those messages arrive:

StepYour app sendsWhat happens on the server
1"Hello from "Text waits in the buffer.
2"sentence boundary streaming."The period may need more text to confirm a sentence boundary.
3" Next we will finish this thought"Hello from sentence boundary streaming. becomes eligible for synthesis. Next we will finish this thought stays buffered.
4closeContextThe remaining text becomes eligible for synthesis, even without final punctuation. The server finishes sending audio, then sends contextClosed.

A confirmed sentence is ready to be synthesized, but may be combined with other ready text while earlier audio is streaming.

Close a context in any mode

In all three modes, send closeContext when you have no more text for this context. It requests synthesis of any remaining buffered text and closes the context after the response finishes. You cannot send more text to that context after requesting close.

javascript
ws.send(JSON.stringify({contextId: "reply-1", closeContext: {}}));

Wait for contextClosed before you stop receiving audio for that context. Sending closeContext begins closing it; audio may still be arriving. Closing the WebSocket immediately can cut off the remaining audio.

Advanced: synthesis batches and timestamps

The server can combine ready sentences while earlier audio is streaming, including when a flush or close releases the remaining text. Do not expect one audio response or flushCompleted event per sendText message or detected sentence.

flushCompleted marks the end of a synthesis batch (a flush). Alignment timestamps reset for the next flush.

Markup and unfinished input

With SENTENCE_BOUNDARY, a supported markup tag can span multiple sendText messages:

text
First message:   "[whis"
Second message:  "per] Hello."
Combined text:   "[whisper] Hello."
  • An incomplete tag stays buffered; the server does not send a partial tag to synthesis.
  • Instructions stay with the speech that follows them. See Voice steering for model-specific instruction support.
  • Finish every markup tag before flushing or closing. Incomplete markup at end of input returns INVALID_ARGUMENT.
  • Ordinary text does not need final punctuation: an explicit flush or close releases it.

A complete break tag, such as <break time="300ms"/>, provides a boundary and produces the requested pause.

Text limits

LimitWhat it means for your app
2,000 characters per sendText messageSplit longer input across multiple messages.
UTF-16 character countingRequest limits and billed usage count UTF-16 code units. Most characters count as one; many emoji count as two.
Buffering safeguardsThe server can flush accumulated text before your app requests it. In SENTENCE_BOUNDARY mode, safeguards can force a cut while keeping supported markup intact. Leave them at their defaults for normal use.
No buffer delay timer in auto modeExplicitly flush or close when your text source finishes so short remaining text is not left waiting.

For optional bufferCharThreshold and maxBufferDelayMs settings, see the API reference.

Have one large block of text? The Streaming API accepts up to 4,000 characters per request. For connection and concurrency limits, see Rate limits.

Timestamp transport strategy

When using timestamp alignment, you can choose how timestamps are delivered alongside audio using timestampTransportStrategy:

  • SYNC (default): Each chunk contains both audio and timestamps together.
  • ASYNC: Audio chunks arrive first, with timestamps following in separate trailing messages. This reduces time-to-first-audio.

See Timestamps for details on how each mode works.

Code examples

API reference

Synthesize Speech WebSocket

View the complete API specification

Next steps