Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Realtime API overview

Build low-latency voice agents over WebSocket or WebRTC, with turn detection and interruptions handled for you.

The Realtime API connects STT, the LLM you choose, and TTS in one spoken conversation. It uses one WebSocket or WebRTC connection and handles turn detection and interruptions. It follows the OpenAI Realtime protocol, so you can migrate an OpenAI Realtime client with few changes. Inworld extensions add back-channel responses, fillers, and memory.

Using AI to code? Give your assistant the docs index at https://docs.inworld.ai/llms.txt. For live search, add the MCP server.

Prefer the terminal? Install the Inworld CLI with npm install -g @inworld/cli. Use it to create and manage API keys for the Realtime API. AI agents can use it too.

Your first session

Connect and set up a session, then stream audio or text and handle the responses:

text
wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime
  1. Connect with Authorization: Basic <api-key> from a server, or a session token from a browser. You receive session.created.
  2. Send session.update with your instructions, model, voice, and tools.
  3. Stream audio with input_audio_buffer.append, or text with conversation.item.create.
  4. Handle response.output_* events until response.done.

Explore

For voice output without speech input, use Voice responses. Generate and speak an answer in one HTTP request.