> ## Documentation Index
> Fetch the complete documentation index at: https://docs.inworld.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Stream chat completions from LLM Router as server-sent events: chunk format, usage, tool calls, and how errors are reported mid-stream.

Set `stream: true` on a chat completion to receive the response as it is generated, instead of waiting for the full message. The stream uses server-sent events (SSE) in the same format as OpenAI's chat completions API, for every provider.

## Stream a response

<CodeGroup>
```bash cURL
curl --no-buffer --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "openai/gpt-5",
    "messages": [{"role": "user", "content": "Write a haiku about routing."}],
    "stream": true
  }'
```

```python Python
import os
from openai import OpenAI

api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
    base_url="https://api.inworld.ai/v1",
    api_key=api_key,
    default_headers={"Authorization": f"Basic {api_key}"},
)

stream = client.chat.completions.create(
    model="openai/gpt-5",
    messages=[{"role": "user", "content": "Write a haiku about routing."}],
    stream=True,
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```
</CodeGroup>

## Stream format

Each event is a `data:` line with a JSON chunk, followed by a blank line. The stream ends with `data: [DONE]`.

```text
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":null},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{"content":"Requests"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","model":"openai/gpt-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

| Field | Description |
|-------|-------------|
| `id` | The same `chatcmpl-` identifier on every chunk of a response. |
| `object` | Always `chat.completion.chunk`. |
| `model` | The provider and model that served the request, which can differ from the one you asked for after routing or a fallback. |
| `choices[].delta` | The new content since the last chunk: `content`, `tool_calls`, or `reasoning`. |
| `choices[].finish_reason` | `null` until the last chunk, then `stop`, `length`, `tool_calls`, or `content_filter`. |
| `metadata` | On the first chunk only: the request ID and the routing [attempts](https://docs.inworld.ai/router/core-concepts/overview.md#fallbacks) that led to this model. |

## Usage

Token usage is included in the stream whenever the provider reports it. You do not need `stream_options`, and the router ignores it.

The chunk that carries `usage` also has a `choices` entry with an empty `delta`, rather than an empty `choices` array as in OpenAI's API. Read `usage` from whichever chunk includes it:

```python Python
for chunk in stream:
    if chunk.usage:
        print(chunk.usage.prompt_tokens, chunk.usage.completion_tokens)
```

## Tool calls and reasoning

- **Tool calls** arrive as `delta.tool_calls` fragments that you assemble by `index`. See [Tool calling](https://docs.inworld.ai/router/capabilities/tool-calling.md#streaming).
- **Reasoning** text arrives in `delta.reasoning` for models that return it. See [Reasoning](https://docs.inworld.ai/router/capabilities/reasoning.md).

## Errors

How an error reaches you depends on when it happens:

- **Before the first chunk**: you get a normal JSON error response with its HTTP status code, not a stream. [Fallbacks](https://docs.inworld.ai/router/core-concepts/overview.md#fallbacks) run at this stage, so you only see an error if every model failed.
- **After the stream has started**: the HTTP status is already `200`. The router sends one `data: {"error": {"message": "..."}}` event and closes the stream without `data: [DONE]`. The router does not fall back to another model once the first chunk has been sent.

Treat a stream that contains an `error` event, or that ends without `[DONE]`, as a failed request.

## Streaming on other APIs

- [Responses API](https://docs.inworld.ai/router/responses.md#streaming) streams typed events in the OpenAI Responses format.
- [Anthropic Messages API](https://docs.inworld.ai/router/anthropic-compatibility.md) streams in Anthropic's event format.
- [Voice responses](https://docs.inworld.ai/router/guides/llm-plus-tts.md#streaming-response-format) add audio chunks to the chat completions stream.

## Next steps

<CardGroup cols={2}>
  <Card title="Tool calling" icon="plug" href="https://docs.inworld.ai/router/capabilities/tool-calling.md">
    Assemble streamed tool calls and return results.
  </Card>

  <Card title="API reference" icon="book" href="https://docs.inworld.ai/api-reference/routerAPI/chat-completions.md">
    Every chat completions parameter.
  </Card>
</CardGroup>
