Capabilities
Reasoning
Control how much a model reasons before it answers with reasoning_effort on LLM Router, and read reasoning text and token usage in the response.
Reasoning models think through a problem before they answer. On LLM Router you control how much reasoning a model does with one parameter, reasoning_effort, and the router translates it to each provider's own control: reasoning effort on OpenAI, thinking on Anthropic, and thinking configuration on Google.
Set the reasoning effort
curl --request POST \
--url https://api.inworld.ai/v1/chat/completions \
--header "Authorization: Basic $INWORLD_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "openai/gpt-5",
"messages": [
{"role": "user", "content": "A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?"}
],
"reasoning_effort": "high"
}'import os
from openai import OpenAI
api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
base_url="https://api.inworld.ai/v1",
api_key=api_key,
default_headers={"Authorization": f"Basic {api_key}"},
)
response = client.chat.completions.create(
model="openai/gpt-5",
messages=[
{"role": "user", "content": "A bat and a ball cost $1.10. The bat costs $1 more than the ball. How much is the ball?"}
],
reasoning_effort="high",
)reasoning_effort accepts none, minimal, low, medium, high, xhigh, and max, in any letter case. Higher values give the model more room to reason, which costs more output tokens and takes longer. none turns reasoning off on models that allow it.
Each model supports its own range of levels. On OpenAI, Anthropic and Inworld-hosted models, a level the model does not offer is mapped to one it does rather than rejected. An unrecognized value is ignored.
On Inworld-hosted models, reasoning is off unless you ask for it.
Set a token budget
To cap reasoning in tokens instead of by level, send one of these. The router accepts them at the top level of the request or inside extra_body:
| Parameter | Example |
|---|---|
reasoning | {"effort": "high", "max_tokens": 4000} |
thinking | {"type": "enabled", "budget_tokens": 4000} |
max_thinking_tokens | 4000 |
With the OpenAI SDK, pass them through extra_body:
response = client.chat.completions.create(
model="anthropic/claude-sonnet-4-6",
messages=[{"role": "user", "content": "Plan a three-step migration."}],
extra_body={"reasoning": {"max_tokens": 4000}},
)When more than one is present:
reasoning_efforttakes precedence overreasoning.effort.- For the budget,
reasoning.max_tokenstakes precedence overthinking.budget_tokens, which takes precedence overmax_thinking_tokens. - A top-level parameter takes precedence over the same parameter in
extra_body.
You can combine an effort level with a budget. Anthropic models require a budget of at least 1,024 tokens.
Models that do not support reasoning
If a model does not support reasoning effort, the router drops the parameter and still serves the request. The dropped parameter is reported as a warning on that attempt in metadata.attempts[].warnings, so you can tell that it had no effect.
Read the reasoning
When the provider returns reasoning text, it is in message.reasoning, next to message.content:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "The ball costs $0.05.",
"reasoning": "Let the ball cost x. Then the bat costs x + 1.00..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 32,
"completion_tokens": 214,
"total_tokens": 246,
"completion_tokens_details": {"reasoning_tokens": 192}
}
}- With
stream: true, reasoning text arrives indelta.reasoning. See Streaming. - Reasoning tokens are counted in
usage.completion_tokens_details.reasoning_tokens, which is present when the count is above zero. - Not every provider returns reasoning text. A model can reason, and report reasoning tokens, without exposing what it reasoned.
The router does not accept a reasoning field on assistant messages you send back, so reasoning from an earlier turn is not carried into the next request.
Reasoning on other APIs
- Responses API: the
reasoningparameter. - Anthropic Messages API: the
thinkingparameter. - Realtime API: reasoning effort for a realtime session.