Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Configuring models

The Inworld Realtime API uses an OpenAI Realtime API-compatible event system to facilitate voice experiences. This guide walks through configuring each layer — STT, LLM, and TTS — plus the conversation- and observability-level controls that span all three. For the field-by-field reference of Inworld extensions, see Inworld Realtime API Extensions.

Configure a session

For WebSocket, the connection starts with a session.created event. For WebRTC, send session.update as soon as the data channel opens. In both cases, use session.update to configure your session. Here you can set:

  • model — LLM provider and model (e.g. openai/gpt-4.1-nano) or router (e.g. inworld/latency-optimizer-ab-test)
  • instructions
  • output_modalities (["audio", "text"], ["audio"], or ["text"])
  • Audio input and output configuration — voice, TTS model, PCM format, speed
  • max_output_tokens ("inf" or a numeric ceiling)
  • tools (function definitions) and tool_choice settings
  • providerData — Inworld extensions for STT, TTS, memory, back-channel, and responsiveness (see Inworld Realtime API Extensions)

Partial updates are supported, so you can adjust the LLM, voice, TTS model, temperature, or tool lists mid-session without rebuilding the socket.

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    type: 'realtime',
    model: 'openai/gpt-4o-mini',
    instructions: 'You are a friendly narrator.',
    output_modalities: ['audio', 'text'],
    temperature: 0.8,
    audio: {
      input: {
        transcription: { model: 'inworld/inworld-stt-1' },
        turn_detection: {
          type: 'semantic_vad',
          create_response: true,
          interrupt_response: true
        }
      },
      output: {
        voice: 'Clive',
        model: 'inworld-tts-2',
        speed: 1.0
      }
    }
  }
}));

STT (Speech-to-Text)

Choose an STT model

Set audio.input.transcription.model to select the speech-to-text model used to transcribe user audio. inworld/inworld-stt-1 is the default for realtime voice agents.

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: { input: { transcription: { model: 'inworld/inworld-stt-1' } } }
  }
}));
ModelBest for
inworld/inworld-stt-1Inworld's first-party STT with configurable turn-taking controls.

If the selected model is not recognised, the server responds with an error event (type: "invalid_request_error", code: "invalid_value", param: "session.audio.input.transcription.model") and the rest of the session.update is not applied. See STT Introduction for the full model catalogue.

Transcription hints

Guide Inworld STT with expected words or phrases in audio.input.transcription.prompts. Use an array of terms, such as names or domain-specific vocabulary. The singular audio.input.transcription.prompt and providerData.stt.prompt fields are not supported by Inworld STT; use prompts instead:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: {
      input: {
        transcription: {
          model: 'inworld/inworld-stt-1',
          prompts: ['angioplasty', 'myocardial infarction'],
          language: 'en'
        }
      }
    }
  }
}));
FieldTypeDescription
modelstringSTT model ID. See Choose an STT model.
promptsstring[]Expected words or phrases to help Inworld STT recognize names and domain-specific vocabulary.
languagestringBCP-47 language code (e.g. "en", "es"). Optional; the model auto-detects when omitted.

Tune turn detection

Turn detection — when the server decides a user has finished speaking — is controlled by the OpenAI-standard audio.input.turn_detection object. The Realtime API supports both VAD types and is wire-compatible with the OpenAI SDK.

semantic_vad

Model-based end-of-turn detection backed by the STT stream. eagerness is the primary tuning knob.

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: {
      input: {
        turn_detection: {
          type: 'semantic_vad',
          eagerness: 'medium',          // 'low' | 'medium' | 'high'
          create_response: true,
          interrupt_response: true
        }
      }
    }
  }
}));
FieldTypeDescription
typestring"semantic_vad" (default)
eagernessstringHow aggressively to end turns: "low", "medium", "high". Lower eagerness requires stronger end-of-turn confidence; higher eagerness commits to end-of-turn sooner.
create_responsebooleanAuto-create a response on turn end (default true)
interrupt_responsebooleanInterrupt the active response when the user speaks (default true)

eagerness maps to a full set of four STT turn-detection parameters — confidence threshold, VAD threshold, minimum end-of-turn silence, and maximum within-turn silence. Higher eagerness uses a lower end-of-turn confidence threshold and shorter silence thresholds to finish turns sooner. The VAD threshold controls which audio is treated as speech.

eagernessend_of_turn_confidence_thresholdvad_thresholdmin_end_of_turn_silence (ms)max_turn_silence (ms)
low0.900.153201500
medium0.750.25160900
high0.650.3030450

These presets are tuned for inworld/inworld-stt-1 and may behave differently if you migrated from a different STT model. Test turn boundaries with your application's audio when migrating.

For inworld/inworld-stt-1, minimum silence determines when the server starts checking for turn completion. Once that minimum is reached, sufficient end-of-turn confidence can finish the turn; maximum silence provides a fallback when confidence is insufficient. These are silence thresholds, not guarantees of transcript delivery or response latency.

Any explicit field under providerData.stt (see STT extensions below) overrides the eagerness-derived default for that field — fields you do not set keep the eagerness mapping.

server_vad

Inworld-hosted Silero VAD + Smart Turn detector. Tunable fields match OpenAI's server_vad shape and can be changed mid-session via partial session.update.

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: {
      input: {
        turn_detection: {
          type: 'server_vad',
          threshold: 0.5,
          prefix_padding_ms: 200,
          silence_duration_ms: 1000,
          idle_timeout_ms: 8000,
          create_response: true,
          interrupt_response: true
        }
      }
    }
  }
}));
FieldTypeDescription
typestring"server_vad"
thresholdnumberSilero VAD speech cutoff, 0.0–1.0. Default 0.5.
prefix_padding_msintegerPre-speech audio retained before an utterance, in ms. Default 200.
silence_duration_msintegerTrailing silence required to finalize the turn, in ms. Default 1000.
idle_timeout_msinteger | nullWhen set, the server emits input_audio_buffer.timeout_triggered after this many ms with no detected speech. null or 0 disables.
create_responsebooleanAuto-create a response on turn end (default true)
interrupt_responsebooleanInterrupt the active response when the user speaks (default true)

All fields accept partial session.update — omit a field to keep its current value. Changes take effect on the next audio chunk processed.

See Turn detection for the speech event lifecycle.

Audio input formats

Set the wire format for client → server audio under audio.input.format. Four formats are supported; pick based on your source. The same catalog applies to audio.output.format (covered under TTS below).

typeEncodingSample rateWhen to use
audio/pcmSigned 16-bit little-endian PCMrate (default 24000)Default for browser, mobile, and most server-side sources. Send mono.
audio/pcmuG.711 μ-lawFixed 8000 Hz (ignore rate)Telephony (Twilio Media Streams, SIP trunks in North America/Japan).
audio/pcmaG.711 A-lawFixed 8000 Hz (ignore rate)Telephony (SIP trunks in Europe and most of the rest of the world).
audio/float3232-bit float PCM, little-endianrate (default 24000)Pipelines that natively produce float32 samples (some audio frameworks).

Audio is always mono, base64-encoded inside the JSON envelope (e.g. input_audio_buffer.append).

javascript
// PCM16 @ 24 kHz (default — omit `format` entirely for the same result)
audio: { input: { format: { type: 'audio/pcm', rate: 24000 } } }

// G.711 μ-law @ 8 kHz, for Twilio
audio: { input: { format: { type: 'audio/pcmu' } } }

A legacy shorthand is also accepted: send format as a bare string — "pcm16", "g711_ulaw", "g711_alaw", or "float32" — and the server expands it to the object form above.

The server resamples to 16 kHz internally for STT, so PCM input rate doesn't need to match the STT model's native rate. Send any rate that's convenient; 24000 and 8000 (G.711) are the common choices.

See Telephony with Twilio for a worked example of the G.711 path.

Send audio input

There are two ways to send audio input:

Method 1: Streaming Audio (Real-time) Use input_audio_buffer.* events for streaming real-time audio from a microphone:

  1. Encode microphone data in your chosen input format (PCM16 at 24 kHz is the default).
  2. Send chunks via input_audio_buffer.append.
  3. VAD automatically detects speech boundaries and commits the buffer.

Method 2: Pre-recorded Audio Use conversation.item.create with input_audio content type for pre-recorded audio chunks:

javascript
ws.send(JSON.stringify({
  type: 'conversation.item.create',
  item: {
    type: 'message',
    role: 'user',
    content: [{
      type: 'input_audio',
      audio: base64AudioData  // Base64-encoded PCM16 or OPUS
    }]
  }
}));

STT extensions

Inworld extensions for STT live under providerData.stt — voice profile signals, language hints, and explicit overrides for the four turn-detection parameters that semantic_vad.eagerness controls implicitly. Full field reference and the voice-profile payload shape are in providerData.stt.


LLM

Choose a router or LLM

Set model in session.update to select which Router or LLM handles the conversation. The format is provider/modelName or inworld/routerId:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    model: 'openai/gpt-4o-mini'
  }
}));

If you omit model, the default model (google-ai-studio/gemini-2.5-flash) is used. You can change the model mid-session with a partial update — the new model takes effect on the next response.

Reasoning effort

For models that support chain-of-thought reasoning (e.g. google-ai-studio/gemini-2.5-pro), configure reasoning depth via text_generation_config.reasoning:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    model: 'google-ai-studio/gemini-2.5-pro',
    text_generation_config: {
      reasoning: { effort: 'MEDIUM' }
    }
  }
}));
EffortBehaviour
NONEDisables reasoning entirely.
MINIMAL~10% of max completion tokens used for reasoning.
LOW~20%
MEDIUM~50% (server default when reasoning is present but effort is omitted).
HIGH~80%
XHIGH~95%

Reasoning tokens are not included in the streamed text output by default; set text_generation_config.reasoning.exclude: false to include them. Usage is reported in response.done under usage.output_token_details.reasoning_tokens.

Support varies by model — some do not support reasoning, others accept only a subset of effort levels. When reasoning is omitted, the model's default applies; for reasoning-capable models this may add latency. Set effort: "NONE" explicitly if you need minimal latency.

For the full field reference (maxTokens, exclude, and other generation params), see text_generation_config.

Send text input

Create explicit conversation items for text turns:

javascript
ws.send(JSON.stringify({
  type: 'conversation.item.create',
  item: {
    type: 'message',
    role: 'user',
    content: [{
      type: 'input_text',
      text: 'Give me a two-sentence summary.'
    }]
  }
}));

Tool calling

Register tools in session.tools so your agent can fetch live data or trigger actions mid-conversation. See Tool calling for registering tools, handling tool calls, and controlling how the conversation continues.

Response metadata

Attach client correlation data to an individual response with response.metadata:

javascript
ws.send(JSON.stringify({
  type: 'response.create',
  response: {
    metadata: {
      request_id: 'req_123',
      workflow: 'horoscope'
    }
  }
}));

The same string map is echoed on response.created and response.done. It is separate from router-facing providerData.metadata. Metadata supports up to 16 entries; keys may contain up to 64 Unicode code points and values up to 512.

Speak exact text without the LLM

response.speak is an Inworld extension that skips the LLM: the server speaks the text you supply, word for word, with the session's TTS. Use it for fixed greetings, scripted prompts, announcements, and farewells. There is no LLM call, so it costs less than response.create and has no wait for the LLM's first token. The trade-off is that nothing varies the wording. The text is spoken verbatim, so any variation has to come from your application.

javascript
ws.send(JSON.stringify({
  type: 'response.speak',
  event_id: 'evt_greeting_1',
  text: 'Hi! Thanks for calling. How can I help you today?',
  force_uninterruptible: true,       // optional, default false
  metadata: { purpose: 'greeting' }  // optional
}));
FieldRequiredDescription
textYesThe text to speak, as a string. It goes to TTS unchanged. An empty string is accepted, and there is no length limit specific to this event.
event_idNoEchoed as request_event_id on the acknowledgment. It is not an idempotency key: sending the same request twice speaks twice.
force_uninterruptibleNoDefault false. When true, protects this response from automatic interruption. See Interruption and active responses.
metadataNoString map with the same limits as response.metadata. Echoed on the acknowledgment, response.created, and response.done.

The event has no output_modalities, voice, instructions, or tools fields. Output follows the session's output_modalities. Audio uses the session's audio.output voice, model, speed, and format, plus its providerData.tts settings, as they were when the request was accepted. A later session.update affects only later responses.

Events you receive

The server first answers with response.speak.acknowledged:

json
{
  "type": "response.speak.acknowledged",
  "event_id": "evt_server_1",
  "request_event_id": "evt_greeting_1",
  "accepted": true,
  "response_id": "resp_1",
  "metadata": { "purpose": "greeting" }
}
  • Accepted: the acknowledgment carries response_id and always arrives before that response's response.created. The usual response events follow: response.output_item.added and response.content_part.added, then response.output_audio.delta and response.output_audio_transcript.delta in audio sessions, or a single response.output_text.delta in text-only sessions. The matching *.done events and one response.done end the response. Acceptance means the request was admitted, not that synthesis succeeded. A TTS failure after acceptance arrives as an error event followed by a response.done with status failed. It does not fall back to the LLM or to text-only output.
  • Rejected: accepted is false and an error object names the problem, for example missing_required_parameter with param: "text", or invalid_value with param: "metadata". There is no response_id, no response is created, and any active response keeps running. A payload that does not decode at all, such as a non-string text, gets an error event with code invalid_json over WebSocket instead of an acknowledgment.

In audio sessions, the transcript follows the audio that has actually been synthesized. With TTS timestamp alignment enabled, transcript deltas arrive as aligned words are synthesized. Without alignment, the transcript arrives in one delta after the audio.

Conversation history

The spoken text becomes an ordinary assistant message item in the conversation, and later responses see it as context. No user item is created, and the input audio buffer is not touched. If synthesis fails or is cancelled partway, the item is stored as incomplete and keeps at most the text that was synthesized. As with any assistant audio, you can send conversation.item.truncate after your client stops playback, including after response.done.

Interruption and active responses

response.speak shares the active-response slot with response.create. A new response.speak or response.create replaces whichever response is active, and response.cancel cancels a spoken response as usual. Without force_uninterruptible, a spoken response is interrupted in the same way as one from response.create.

With force_uninterruptible: true:

  • Detected user speech and new user messages do not cancel the response. User input is still processed and stored.
  • Automatically triggered responses, such as a turn-detection response or an automatic tool continuation, are rejected with an error event whose code is conversation_already_has_active_response. They are not replayed later. Send response.create when you want the agent to continue.
  • Explicit response.cancel, response.create, and response.speak still cancel or replace the response.
  • Protection ends when generation finishes, not when your client finishes playing the audio. Keep your own player from being interrupted until buffered audio has played. The assistant item carries force_uninterruptible: true, but it does not protect later responses.

Usage

No LLM is called, so response.done reports zero tokens and no usage.llm block. In audio sessions, usage.tts reports the synthesis: characters counts the whole supplied text, and audio_seconds the audio generated. Text-only sessions don't open TTS and report no usage.tts. Spoken responses don't trigger memory summarization or responsiveness fillers.

Memory

Inworld's automatic conversation memory layer extracts durable facts and a rolling summary, prepends them to the system prompt, and trims older transcript items so context stays bounded. Configured under providerData.memory. See providerData.memory for the field reference, and Long-term Memory for the cross-session persistence pattern.


TTS (Text-to-Speech)

Choose a TTS model

Set audio.output.model to select the text-to-speech model:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: { output: { model: 'inworld-tts-2' } }
  }
}));
ModelNotes
inworld-tts-2Best quality, steerability, and multilingual coverage. Recommended for most agents; server default when audio.output.model is omitted. Required for providerData.tts.conversational and the CREATIVE delivery mode.
inworld-tts-2-flashOur fastest, most cost-efficient model — same language coverage as inworld-tts-2. Does not support steering instructions (non-verbal tags like [laugh] still work).

Examples throughout these docs use inworld-tts-2, which is the server default and the right choice for production agents — it leads on quality and concurrency. Switch to inworld-tts-2-flash when latency or cost per character is the deciding factor. You can change the TTS model mid-session alongside voice or independently.

Choose a voice

Set audio.output.voice to control the agent's speaking voice:

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    audio: { output: { voice: 'Olivia' } }
  }
}));

The default voice is Dennis. Browse available voices in the TTS Playground or list them programmatically with the List Voices API.

Audio output format

Set the wire format for server → client audio under audio.output.format. The catalog is identical to Audio input formats above — audio/pcm, audio/pcmu, audio/pcma, or audio/float32. Default is PCM16 at 24 kHz.

javascript
// Default — PCM16 @ 24 kHz (omit format entirely for the same result)
audio: { output: { format: { type: 'audio/pcm', rate: 24000 } } }

// G.711 μ-law @ 8 kHz, for Twilio — TTS audio comes back already mulaw-encoded
audio: { output: { format: { type: 'audio/pcmu' } } }

In most setups, set input and output to the same format so your client only has one codec path. The exception is telephony, where you typically want both sides on G.711 to match the carrier.

The server resamples internally — TTS models synthesize at their native rate and the server downsamples (or upsamples) to whatever audio.output.format.rate you request, so any reasonable rate is accepted.

TTS extensions

Inworld extensions for TTS live under providerData.tts — segmentation strategy, steering handling, synthesis language, the TTS-2 delivery preset, (for TTS-2) conversational mode, and timestamp alignment for lip-sync or word highlighting. Full field reference, segmenter strategy table, conversational-mode details, and the timestamp output shape are in providerData.tts.

To opt into alignment, set providerData.tts.timestamp_type to WORD or CHARACTER and choose a transport strategy (SYNC for real-time lip-sync, ASYNC for lower latency). See TTS timestamps and alignment for the full output shape and sync/async semantics.


Managing the session

Conversation state

Use conversation events to keep context lean:

  • conversation.item.retrieve: pull any prior item by ID.
  • conversation.item.delete: remove items that should not remain in context.

Pair these with max_output_tokens and response.cancel to control overall cost (conversation management guide).

Observing usage

response.done carries a response.usage block on every response — including cancelled responses (interruption, supersede). The base fields (total_tokens, input_tokens, output_tokens, plus input_token_details / output_token_details) cover LLM accounting, and three optional sub-objects attribute usage per modality:

FieldTypeDescription
usage.llm.modelstringEffective upstream LLM after router resolution. Useful when you sent inworld/auto and want to see which model the router picked.
usage.tts.modelstringTTS model used (e.g. inworld-tts-2).
usage.tts.charactersintegerCharacters synthesized across all TTS segments of this response.
usage.tts.audio_secondsnumberAssistant audio duration emitted by TTS, in seconds. The canonical TTS billing signal.
usage.stt.modelstringSTT model used (e.g. inworld/inworld-stt-1).
usage.stt.audio_secondsnumberUser audio duration transcribed for this turn, in seconds. Drained per response.done from a rolling per-session counter — each response sees only the user audio that arrived since the previous response.done.

Each modality sub-object is omitted when there's nothing to report (e.g. a TTS-only response with no preceding user turn won't carry stt).

javascript
ws.on('message', (buffer) => {
  const event = JSON.parse(buffer.toString());
  if (event.type !== 'response.done') return;

  const u = event.response.usage;
  if (!u) return;
  console.log(
    `[usage] llm=${u.llm?.model ?? '?'} ` +
    `tokens=${u.input_tokens}/${u.output_tokens} ` +
    `tts=${u.tts?.audio_seconds?.toFixed(2) ?? '0'}s ` +
    `stt=${u.stt?.audio_seconds?.toFixed(2) ?? '0'}s`
  );
});

For the full schema including input_token_details / output_token_details breakdowns, see the response.done event.

input_token_details also carries prompt-cache counters: cached_tokens (input served from a cache hit) and cache_write_tokens (input written when establishing a cache entry). These appear automatically when a provider caches implicitly, and you can opt into explicit caching of the system prompt and tools via providerData.caching.

Monitor errors

Handle error events (with type, code, and param) and implement a reconnection/backoff strategy for transient failures. See the API reference for error event schemas.

javascript
ws.on('message', (buffer) => {
  const event = JSON.parse(buffer.toString());

  if (event.type === 'error') {
    handleError(event.error);
  }
});