Realtime API overview
Build low-latency voice agents over WebSocket or WebRTC, with turn detection and interruptions handled for you.
The Realtime API connects STT, the LLM you choose, and TTS in one spoken conversation. It uses one WebSocket or WebRTC connection and handles turn detection and interruptions. It follows the OpenAI Realtime protocol, so you can migrate an OpenAI Realtime client with few changes. Inworld extensions add back-channel responses, fillers, and memory.
WebSocket quickstart
A Node.js server and a browser client, streaming audio both ways.
WebRTC quickstart
Voice using the browser's built-in support, with automatic echo cancellation.
Realtime Playground
Talk to a voice agent in the Portal.
JS examples
Complete Node.js and browser projects.
Python examples
Complete Python projects.
API reference
Every event, for WebSocket and WebRTC.
Using AI to code? Give your assistant the docs index at https://docs.inworld.ai/llms.txt. For live search, add the MCP server.
Prefer the terminal? Install the Inworld CLI with npm install -g @inworld/cli. Use it to create and manage API keys for the Realtime API. AI agents can use it too.
Your first session
Connect and set up a session, then stream audio or text and handle the responses:
wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime- Connect with
Authorization: Basic <api-key>from a server, or a session token from a browser. You receivesession.created. - Send
session.updatewith your instructions, model, voice, and tools. - Stream audio with
input_audio_buffer.append, or text withconversation.item.create. - Handle
response.output_*events untilresponse.done.
Explore
Configuring models
Choose the STT model, the LLM or router, the voice, and turn detection.
Tool calling
Let the agent call your functions mid-conversation.
Managing conversations
Conversation items, turn detection events, interruptions, and errors.
Adding naturalness
Steering, back-channel responses, and fillers that keep the agent feeling present.
Twilio integration
Put the agent on a phone number.
For voice output without speech input, use Voice responses. Generate and speak an answer in one HTTP request.