Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Streaming

Stream chat completions from LLM Router as server-sent events: chunk format, usage, tool calls, and how errors are reported mid-stream.

Set stream: true on a chat completion to receive the response as it is generated, instead of waiting for the full message. The stream uses server-sent events (SSE) in the same format as OpenAI's chat completions API, for every provider.

Stream a response

cURL
curl --no-buffer --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "openai/gpt-5",
    "messages": [{"role": "user", "content": "Write a haiku about routing."}],
    "stream": true
  }'

Stream format

Each event is a data: line with a JSON chunk, followed by a blank line. The stream ends with data: [DONE].

text
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":null},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"content":"Requests"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
FieldDescription
idThe same chatcmpl- identifier on every chunk of a response.
objectAlways chat.completion.chunk.
modelThe provider and model that served the request, which can differ from the one you asked for after routing or a fallback.
choices[].deltaThe new content since the last chunk: content, tool_calls, or reasoning.
choices[].finish_reasonnull until the last chunk, then stop, length, tool_calls, or content_filter.
metadataOn the first chunk only: the request ID and the routing attempts that led to this model.

Usage

Token usage is included in the stream whenever the provider reports it. You do not need stream_options, and the router ignores it.

The chunk that carries usage also has a choices entry with an empty delta, rather than an empty choices array as in OpenAI's API. Read usage from whichever chunk includes it:

Python
for chunk in stream:
    if chunk.usage:
        print(chunk.usage.prompt_tokens, chunk.usage.completion_tokens)

Tool calls and reasoning

  • Tool calls arrive as delta.tool_calls fragments that you assemble by index. See Tool calling.
  • Reasoning text arrives in delta.reasoning for models that return it. See Reasoning.

Errors

How an error reaches you depends on when it happens:

  • Before the first chunk: you get a normal JSON error response with its HTTP status code, not a stream. Fallbacks run at this stage, so you only see an error if every model failed.
  • After the stream has started: the HTTP status is already 200. The router sends one data: {"error": {"message": "..."}} event and closes the stream without data: [DONE]. The router does not fall back to another model once the first chunk has been sent.

Treat a stream that contains an error event, or that ends without [DONE], as a failed request.

Streaming on other APIs

Next steps