Skip to main content
The Inworld Realtime API uses an OpenAI Realtime API-compatible event system to facilitate voice experiences. This guide walks through configuring each layer — STT, LLM, and TTS — plus the conversation- and observability-level controls that span all three. For the field-by-field reference of Inworld extensions, see Inworld Realtime API Extensions.

Configure a session

For WebSocket, the connection starts with a session.created event. For WebRTC, send session.update as soon as the data channel opens. In both cases, use session.update to configure your session. Here you can set:
  • model — LLM provider and model (e.g. openai/gpt-4.1-nano) or router (e.g. inworld/latency-optimizer-ab-test)
  • instructions
  • output_modalities (["audio", "text"], ["audio"], or ["text"])
  • Audio input and output configuration — voice, TTS model, PCM format, speed
  • max_output_tokens ("inf" or a numeric ceiling)
  • tools (function definitions) and tool_choice settings
  • providerData — Inworld extensions for STT, TTS, memory, back-channel, and responsiveness (see Inworld Realtime API Extensions)
Partial updates are supported, so you can adjust the LLM, voice, TTS model, temperature, or tool lists mid-session without rebuilding the socket.

STT (Speech-to-Text)

Choose an STT model

Set audio.input.transcription.model to select the speech-to-text model used to transcribe user audio. inworld/inworld-stt-1 is the recommended default for most realtime voice agents; pick a third-party model when its specific strength (sub-300ms latency, semantic end-of-turn, etc.) matters for your use case.
If the selected model is not recognised, the server responds with an error event (type: "invalid_request_error", code: "invalid_value", param: "session.audio.input.transcription.model") and the rest of the session.update is not applied. See STT Introduction for the full model catalogue and comparison.

Transcription hints

Guide the STT decoder with a prompt (vocabulary, domain context, formatting preferences). This is the OpenAI-standard audio.input.transcription.prompt field and is portable across OpenAI-compatible SDKs:

Tune turn detection

Turn detection — when the server decides a user has finished speaking — is controlled by the OpenAI-standard audio.input.turn_detection object. The Realtime API supports both VAD types and is wire-compatible with the OpenAI SDK.

semantic_vad

Model-based end-of-turn detection backed by the STT stream. eagerness is the primary tuning knob.
eagerness maps to a full set of four STT turn-detection parameters — confidence threshold, VAD threshold, minimum end-of-turn silence, and maximum within-turn silence. Lower thresholds and shorter silences mean the STT model commits to end-of-turn sooner (more eager). auto mirrors medium until router-side adaptive logic exists. Any explicit field under providerData.stt (see STT extensions below) overrides the eagerness-derived default for that field — fields you do not set keep the eagerness mapping.

server_vad

Inworld-hosted Silero VAD + Smart Turn detector. Tunable fields match OpenAI’s server_vad shape and can be changed mid-session via partial session.update.
All fields accept partial session.update — omit a field to keep its current value. Changes take effect on the next audio chunk processed. See Voice Activity Detection (VAD) for the VAD event lifecycle.

Audio input formats

Set the wire format for client → server audio under audio.input.format. Four formats are supported; pick based on your source. The same catalog applies to audio.output.format (covered under TTS below). Audio is always mono, base64-encoded inside the JSON envelope (e.g. input_audio_buffer.append).
A legacy shorthand is also accepted: send format as a bare string — "pcm16", "g711_ulaw", "g711_alaw", or "float32" — and the server expands it to the object form above. The server resamples to 16 kHz internally for STT, so PCM input rate doesn’t need to match the STT model’s native rate. Send any rate that’s convenient; 24000 and 8000 (G.711) are the common choices. See Telephony with Twilio for a worked example of the G.711 path.

Send audio input

There are two ways to send audio input: Method 1: Streaming Audio (Real-time) Use input_audio_buffer.* events for streaming real-time audio from a microphone:
  1. Encode microphone data in your chosen input format (PCM16 at 24 kHz is the default).
  2. Send chunks via input_audio_buffer.append.
  3. VAD automatically detects speech boundaries and commits the buffer.
Method 2: Pre-recorded Audio Use conversation.item.create with input_audio content type for pre-recorded audio chunks:

STT extensions

Inworld extensions for STT live under providerData.stt — voice profile signals, language hints (Soniox), and explicit overrides for the four turn-detection parameters that semantic_vad.eagerness controls implicitly. Full field reference and the voice-profile payload shape are in providerData.stt.

LLM

Choose a router or LLM

Set model in session.update to select which Router or LLM handles the conversation. The format is provider/modelName or inworld/routerId:
If you omit model, the default model (google-ai-studio/gemini-2.5-flash) is used. You can change the model mid-session with a partial update — the new model takes effect on the next response.

Reasoning effort

For models that support chain-of-thought reasoning (e.g. google-ai-studio/gemini-2.5-pro), configure reasoning depth via text_generation_config.reasoning:
Reasoning tokens are not included in the streamed text output by default; set text_generation_config.reasoning.exclude: false to include them. Usage is reported in response.done under usage.output_token_details.reasoning_tokens.
Support varies by model — some do not support reasoning, others accept only a subset of effort levels. When reasoning is omitted, the model’s default applies; for reasoning-capable models this may add latency. Set effort: "NONE" explicitly if you need minimal latency.
For the full field reference (maxTokens, exclude, and other generation params), see text_generation_config.

Send text input

Create explicit conversation items for text turns:

Function calling

The Realtime API supports function calling so your agent can fetch live data or trigger actions mid-conversation. Define functions in session.tools, then handle calls as they arrive.

1. Register a tool

2. Handle the function call

When the model decides to call a function, you receive a response.function_call_arguments.done event with the call_id, function name, and serialized arguments. Execute your logic, then return the result:

3. What happens next

After response.create, the model incorporates the function output and continues the conversation — speaking the horoscope aloud (if output_modalities includes audio) or streaming text deltas. The user hears the answer without any gap in the conversation flow. You can register multiple tools and the model will call them as needed. Each call arrives as a separate response.function_call_arguments.done event with its own call_id.

Memory

Inworld’s automatic conversation memory layer extracts durable facts and a rolling summary, prepends them to the system prompt, and trims older transcript items so context stays bounded. Configured under providerData.memory. See providerData.memory for the field reference, and Long-term Memory for the cross-session persistence pattern.

TTS (Text-to-Speech)

Choose a TTS model

Set audio.output.model to select the text-to-speech model:
Examples throughout these docs use inworld-tts-2 for quality; switch to inworld-tts-1.5-mini if you’re optimizing for raw latency or running at high concurrency. You can change the TTS model mid-session alongside voice or independently.

Choose a voice

Set audio.output.voice to control the agent’s speaking voice:
The default voice is Dennis. Browse available voices in the TTS Playground or list them programmatically with the List Voices API.

Audio output format

Set the wire format for server → client audio under audio.output.format. The catalog is identical to Audio input formats above — audio/pcm, audio/pcmu, audio/pcma, or audio/float32. Default is PCM16 at 24 kHz.
In most setups, set input and output to the same format so your client only has one codec path. The exception is telephony, where you typically want both sides on G.711 to match the carrier. The server resamples internally — TTS models synthesize at their native rate and the server downsamples (or upsamples) to whatever audio.output.format.rate you request, so any reasonable rate is accepted.

TTS extensions

Inworld extensions for TTS live under providerData.tts — segmentation strategy, steering handling, synthesis language, the TTS-2 delivery preset, (for TTS-2) conversational mode, and timestamp alignment for lip-sync or word highlighting. Full field reference, segmenter strategy table, conversational-mode details, and the timestamp output shape are in providerData.tts. To opt into alignment, set providerData.tts.timestamp_type to WORD or CHARACTER and choose a transport strategy (SYNC for real-time lip-sync, ASYNC for lower latency). See TTS timestamps and alignment for the full output shape and sync/async semantics.

Managing the session

Conversation state

Use conversation events to keep context lean:
  • conversation.item.retrieve: pull any prior item by ID.
  • conversation.item.delete: remove items that should not remain in context.
Pair these with max_output_tokens and response.cancel to control overall cost (conversation management guide).

Observing usage

response.done carries a response.usage block on every response — including cancelled responses (barge-in, supersede). The base fields (total_tokens, input_tokens, output_tokens, plus input_token_details / output_token_details) cover LLM accounting, and three optional sub-objects attribute usage per modality: Each modality sub-object is omitted when there’s nothing to report (e.g. a TTS-only response with no preceding user turn won’t carry stt).
For the full schema including input_token_details / output_token_details breakdowns, see the response.done event. input_token_details also carries prompt-cache counters: cached_tokens (input served from a cache hit) and cache_write_tokens (input written when establishing a cache entry). These appear automatically when a provider caches implicitly, and you can opt into explicit caching of the system prompt and tools via providerData.caching.

Monitor errors

Handle error events (with type, code, and param) and implement a reconnection/backoff strategy for transient failures. See the API reference for error event schemas.