Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Connect

WebSocket

Connect to the Realtime API over WebSocket, set up sessions, and stream audio and text events.

Use WebSocket to connect to the Realtime API and exchange audio and text events. For low-latency voice using the browser's built-in support, see WebRTC.

Endpoint

wss://api.inworld.ai/api/v1/realtime/session?key=<session-id>&protocol=realtime
ParameterRequiredDescription
keyYesSession ID from your app
protocolYesrealtime

Authentication

EnvironmentHeaderNotes
Server-side (Node.js)Authorization: Basic <api-key>Your API key from Inworld Portal
Client-side (browser)Authorization: Bearer <jwt-token>Mint a session token on your backend. See the JWT sample app for a complete example

Flow

  1. Connect → receive session.created
  2. Send session.update (instructions, audio config, tools)
  3. Stream audio (input_audio_buffer.append) or text (conversation.item.create)
  4. response.create → handle response.output_* until response.done

Session Config

Use session.update to change your system prompt, voice, model, or tools during a conversation. It accepts partial updates.

javascript
ws.send(JSON.stringify({
  type: 'session.update',
  session: {
    type: 'realtime',
    model: 'openai/gpt-4o-mini',
    instructions: 'You are a concise concierge.',
    output_modalities: ['audio', 'text'],
    audio: {
      input: {
        turn_detection: {
          type: 'semantic_vad',
          eagerness: 'medium',
          create_response: true,
          interrupt_response: true
        }
      },
      output: {
        voice: 'Clive',
        model: 'inworld-tts-2',
        speed: 1.0
      }
    },
    tools: [{
      type: 'function',
      name: 'get_weather',
      description: 'Fetch weather for a location',
      parameters: {
        type: 'object',
        properties: { location: { type: 'string' } },
        required: ['location']
      }
    }],
    providerData: {
      stt: {
        voice_profile: true,
        language_hints: ['en-US', 'es-MX'],
        end_of_turn_confidence_threshold: 0.7,
        min_end_of_turn_silence: 200,
        max_turn_silence: 5000,
        vad_threshold: 0.5
      },
      tts: {
        segmenter_strategy: 'sentence',
        steering_handling: 'emit_once',
        language: 'en-US',
        delivery_mode: 'CREATIVE',
        conversational: false
      },
      memory: {
        enabled: true,
        turn_interval: 5,
        max_facts: 50
      },
      backchannel: {
        enabled: true
      },
      responsiveness: {
        enabled: true
      }
    }
  }
}));

providerData adds Inworld options to the OpenAI-compatible session format, including STT tuning, TTS segmentation and steering, and automatic memory. You can change most fields during a conversation with a partial session.update. A few fields are read only when the session opens and ignored afterwards, notably providerData.tts.conversational and providerData.tts.user_turn_mode. See Inworld Realtime API Extensions for each field and whether it can change during a session.

Audio

The default audio format is base64-encoded mono PCM16 at 24 kHz. For telephony, the API also accepts audio/pcmu and audio/pcma (G.711 μ-law / A-law) at 8 kHz. Use audio/float32 for pipelines that produce 32-bit float samples.

Set the input and output formats with audio.input.format and audio.output.format in session.update. See Audio input formats for the full list. We recommend 60-100ms chunks.

javascript
ws.send(JSON.stringify({ type: 'input_audio_buffer.append', audio: base64PcmChunk }));

Use input_audio_buffer.clear to discard unwanted audio.

Text

To send text from your client, use conversation.item.create.

javascript
ws.send(JSON.stringify({
  type: 'conversation.item.create',
  item: {
    type: 'message',
    role: 'user',
    content: [{
      type: 'input_text',
      text: 'Can you summarize the notes I sent?'
    }]
  }
}));

Events

WebSocket events carry speech-to-speech conversations. Your client sends events to the API and handles events from the server.

  • Session: session.created, session.updated
  • Conversation: conversation.item.added/done/retrieved/deleted/truncated, transcription deltas/completions
  • Responses: response.speak.acknowledged (after a response.speak), response.created, response.output_item.added/done, response.output_text.delta/done, response.output_audio.delta/done, response.done
  • Audio/VAD: input_audio_buffer.speech_started, input_audio_buffer.speech_stopped, input_audio_buffer.turn_suggestion, input_audio_buffer.turn_suggestion_revoked, input_audio_buffer.timeout_triggered, response.output_audio_transcript.delta
  • Back-channel: response.backchannel.audio.delta, response.backchannel.audio.done, response.backchannel.skipped
  • Errors: error

See the API reference for all events and their schemas.

Node.js websocket server example

This server-side Node.js example uses the ws library with Basic auth.

javascript
import WebSocket from 'ws';

const sessionId = 'your-session-id';
const credentials = process.env.INWORLD_API_KEY;

const ws = new WebSocket(`wss://api.inworld.ai/api/v1/realtime/session?key=${sessionId}&protocol=realtime`, {
  headers: {
    Authorization: `Basic ${credentials}`
  }
});

ws.on('open', () => {
  console.log('WebSocket connected');
});

ws.on('message', (buffer) => {
  const message = JSON.parse(buffer.toString());

  switch (message.type) {
    case 'session.created':
      console.log('Session created:', message.session.id);
      updateSession();
      break;
    case 'session.updated':
      console.log('Session updated');
      sendMessage('Hello!');
      break;
    case 'conversation.item.added':
      console.log('Conversation item added:', message.item.id);
      break;
    case 'conversation.item.done':
      console.log('Conversation item done');
      createResponse();
      break;
    case 'input_audio_buffer.speech_started':
      console.log('Speech started at', message.audio_start_ms, 'ms');
      break;
    case 'input_audio_buffer.speech_stopped':
      console.log('Speech stopped at', message.audio_end_ms, 'ms');
      break;
    case 'conversation.item.input_audio_transcription.delta':
      console.log('Transcription delta:', message.delta);
      break;
    case 'conversation.item.input_audio_transcription.completed':
      console.log('Transcription complete:', message.transcript);
      break;
    case 'response.created':
      console.log('Response created:', message.response.id);
      break;
    case 'response.output_item.added':
      console.log('Output item added:', message.item.id);
      break;
    case 'response.output_text.delta':
      console.log('Text delta:', message.delta);
      break;
    case 'response.output_audio.delta':
      // Decode and play audio chunk
      const audioBuffer = Buffer.from(message.delta, 'base64');
      playAudio(audioBuffer);
      break;
    case 'response.output_audio_transcript.delta':
      console.log('Audio transcript delta:', message.delta);
      break;
    case 'response.done':
      console.log('Response complete, status:', message.response.status);
      break;
    case 'error':
      console.error('Error:', message.error.message, message.error.code);
      break;
  }
});

function updateSession() {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: {
      type: 'realtime',
      output_modalities: ['text', 'audio'],
      instructions: 'You are a helpful AI assistant.',
      audio: {
        input: {
          turn_detection: {
            type: 'semantic_vad',
            eagerness: 'medium',
            create_response: true,
            interrupt_response: true
          }
        },
        output: {
          voice: 'Clive'
        }
      }
    }
  }));
}

function sendMessage(text) {
  ws.send(JSON.stringify({
    type: 'conversation.item.create',
    item: {
      type: 'message',
      role: 'user',
      content: [{ type: 'input_text', text }]
    }
  }));
}

function createResponse() {
  ws.send(JSON.stringify({
    type: 'response.create',
    response: {
      output_modalities: ['text', 'audio']
    }
  }));
}

function cancelResponse() {
  ws.send(JSON.stringify({ type: 'response.cancel' }));
}

function sendAudioChunk(audioChunk) {
  ws.send(JSON.stringify({
    type: 'input_audio_buffer.append',
    audio: audioChunk // base64-encoded audio data
  }));
}

function clearAudioBuffer() {
  ws.send(JSON.stringify({ type: 'input_audio_buffer.clear' }));
}

See the API reference for full schemas.