Capabilities
Streaming
Stream chat completions from LLM Router as server-sent events: chunk format, usage, tool calls, and how errors are reported mid-stream.
Set stream: true on a chat completion to receive the response as it is generated, instead of waiting for the full message. The stream uses server-sent events (SSE) in the same format as OpenAI's chat completions API, for every provider.
Stream a response
curl --no-buffer --request POST \
--url https://api.inworld.ai/v1/chat/completions \
--header "Authorization: Basic $INWORLD_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "openai/gpt-5",
"messages": [{"role": "user", "content": "Write a haiku about routing."}],
"stream": true
}'import os
from openai import OpenAI
api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
base_url="https://api.inworld.ai/v1",
api_key=api_key,
default_headers={"Authorization": f"Basic {api_key}"},
)
stream = client.chat.completions.create(
model="openai/gpt-5",
messages=[{"role": "user", "content": "Write a haiku about routing."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Stream format
Each event is a data: line with a JSON chunk, followed by a blank line. The stream ends with data: [DONE].
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":null},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"content":"Requests"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]| Field | Description |
|---|---|
id | The same chatcmpl- identifier on every chunk of a response. |
object | Always chat.completion.chunk. |
model | The provider and model that served the request, which can differ from the one you asked for after routing or a fallback. |
choices[].delta | The new content since the last chunk: content, tool_calls, or reasoning. |
choices[].finish_reason | null until the last chunk, then stop, length, tool_calls, or content_filter. |
metadata | On the first chunk only: the request ID and the routing attempts that led to this model. |
Usage
Token usage is included in the stream whenever the provider reports it. You do not need stream_options, and the router ignores it.
The chunk that carries usage also has a choices entry with an empty delta, rather than an empty choices array as in OpenAI's API. Read usage from whichever chunk includes it:
for chunk in stream:
if chunk.usage:
print(chunk.usage.prompt_tokens, chunk.usage.completion_tokens)Tool calls and reasoning
- Tool calls arrive as
delta.tool_callsfragments that you assemble byindex. See Tool calling. - Reasoning text arrives in
delta.reasoningfor models that return it. See Reasoning.
Errors
How an error reaches you depends on when it happens:
- Before the first chunk: you get a normal JSON error response with its HTTP status code, not a stream. Fallbacks run at this stage, so you only see an error if every model failed.
- After the stream has started: the HTTP status is already
200. The router sends onedata: {"error": {"message": "..."}}event and closes the stream withoutdata: [DONE]. The router does not fall back to another model once the first chunk has been sent.
Treat a stream that contains an error event, or that ends without [DONE], as a failed request.
Streaming on other APIs
- Responses API streams typed events in the OpenAI Responses format.
- Anthropic Messages API streams in Anthropic's event format.
- Voice responses add audio chunks to the chat completions stream.