> ## Documentation Index
> Fetch the complete documentation index at: https://docs.inworld.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Reasoning

> Control how much a model reasons before it answers with reasoning_effort on LLM Router, and read reasoning text and token usage in the response.

Reasoning models think through a problem before they answer. On LLM Router you control how much reasoning a model does with one parameter, `reasoning_effort`, and the router translates it to each provider's own control: reasoning effort on OpenAI, thinking on Anthropic, and thinking configuration on Google.

## Set the reasoning effort

<CodeGroup>
```bash cURL
curl --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "openai/gpt-5",
    "messages": [
      {"role": "user", "content": "A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?"}
    ],
    "reasoning_effort": "high"
  }'
```

```python Python
import os
from openai import OpenAI

api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
    base_url="https://api.inworld.ai/v1",
    api_key=api_key,
    default_headers={"Authorization": f"Basic {api_key}"},
)

response = client.chat.completions.create(
    model="openai/gpt-5",
    messages=[
        {"role": "user", "content": "A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?"}
    ],
    reasoning_effort="high",
)
```
</CodeGroup>

`reasoning_effort` accepts `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`, in any letter case. Higher values give the model more room to reason, which costs more output tokens and takes longer. `none` turns reasoning off on models that allow it.

Each model supports its own range of levels. On OpenAI, Anthropic and Inworld-hosted models, a level the model does not offer is mapped to one it does rather than rejected. An unrecognized value is ignored.

<Note>
  On Inworld-hosted models, reasoning is off unless you ask for it.
</Note>

## Set a token budget

To cap reasoning in tokens instead of by level, send one of these. The router accepts them at the top level of the request or inside `extra_body`:

| Parameter | Example |
|-----------|---------|
| `reasoning` | `{"effort": "high", "max_tokens": 4000}` |
| `thinking` | `{"type": "enabled", "budget_tokens": 4000}` |
| `max_thinking_tokens` | `4000` |

With the OpenAI SDK, pass them through `extra_body`:

```python Python
response = client.chat.completions.create(
    model="anthropic/claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Plan a three-step migration."}],
    extra_body={"reasoning": {"max_tokens": 4000}},
)
```

When more than one is present:

- `reasoning_effort` takes precedence over `reasoning.effort`.
- For the budget, `reasoning.max_tokens` takes precedence over `thinking.budget_tokens`, which takes precedence over `max_thinking_tokens`.
- A top-level parameter takes precedence over the same parameter in `extra_body`.

You can combine an effort level with a budget. Anthropic models require a budget of at least 1,024 tokens.

## Models that do not support reasoning

If a model does not support reasoning effort, the router drops the parameter and still serves the request. The dropped parameter is reported as a warning on that attempt in `metadata.attempts[].warnings`, so you can tell that it had no effect.

## Read the reasoning

When the provider returns reasoning text, it is in `message.reasoning`, next to `message.content`:

```json
{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "The ball costs $0.05.",
        "reasoning": "Let the ball cost x. Then the bat costs x + 1.00..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 214,
    "total_tokens": 246,
    "completion_tokens_details": {"reasoning_tokens": 192}
  }
}
```

- With `stream: true`, reasoning text arrives in `delta.reasoning`. See [Streaming](https://docs.inworld.ai/router/capabilities/streaming.md).
- Reasoning tokens are counted in `usage.completion_tokens_details.reasoning_tokens`, which is present when the count is above zero.
- Not every provider returns reasoning text. A model can reason, and report reasoning tokens, without exposing what it reasoned.

The router does not accept a `reasoning` field on assistant messages you send back, so reasoning from an earlier turn is not carried into the next request.

## Reasoning on other APIs

- [Responses API](https://docs.inworld.ai/router/responses.md): the `reasoning` parameter.
- [Anthropic Messages API](https://docs.inworld.ai/router/anthropic-compatibility.md#request-parameters): the `thinking` parameter.
- [Realtime API](https://docs.inworld.ai/realtime/usage/using-realtime-models.md#reasoning-effort): reasoning effort for a realtime session.

## Next steps

<CardGroup cols={2}>
  <Card title="Tool calling" icon="plug" href="https://docs.inworld.ai/router/capabilities/tool-calling.md">
    Let a reasoning model call your functions.
  </Card>

  <Card title="API reference" icon="book" href="https://docs.inworld.ai/api-reference/routerAPI/chat-completions.md">
    Every chat completions parameter.
  </Card>
</CardGroup>
