Routing
Request-level routing
Route requests without creating a router
With the Chat Completions API, you can route requests without creating a router. In a single request, you can call a specific model, add fallbacks, or let the router choose a model for you.
Use request-level routing to:
- Call a specific model through a unified API without setting up a router
- Prototype or benchmark before choosing a router configuration
For conditional routing, A/B testing with weighted variants, or shared prompt templates, set up a router.
Authentication
Run these examples on your server. Set INWORLD_API_KEY to the Base64 credentials copied from Portal or the CLI. Use the value as-is; don't Base64-encode it again. A Standard key can call chat completions. You only need Router Write permission to change router configuration.
The direct HTTP examples use Basic. Keep API keys on your server. Browser and mobile clients should use a token created by your backend or a server proxy.
Direct model call
Set model to a provider/model identifier:
curl -X POST https://api.inworld.ai/v1/chat/completions \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "openai/gpt-5.2",
"messages": [{ "role": "user", "content": "Hello!" }]
}'This calls the specified model through LLM Router's unified API, with no routing logic.
Fallbacks
Add fallback models via extra_body.models. If the primary model fails, the router automatically tries the next model in the list:
curl -X POST https://api.inworld.ai/v1/chat/completions \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "openai/gpt-5.2",
"messages": [{ "role": "user", "content": "Hello!" }],
"extra_body": {
"models": ["anthropic/claude-opus-4-6", "google-ai-studio/gemini-2.5-pro"]
}
}'In this example, the router tries gpt-5.2 first, then Claude Opus, then Gemini Pro. The response's metadata.attempts array shows which models it tried.
Fallback by first token timeout
Set a time to first token (TTFT) timeout to trigger fallback when a model is slow. If the model doesn't return its first token in time, the router cancels the request and tries the next model.
Use this when your application needs a fast response and can try another model instead of waiting.
Set extra_body.fallback.ttft_timeout in your request:
curl -X POST https://api.inworld.ai/v1/chat/completions \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "openai/gpt-5.2",
"messages": [{ "role": "user", "content": "Hello" }],
"extra_body": {
"models": ["openai/gpt-4o", "google-ai-studio/gemini-2.5-pro"],
"fallback": {
"ttft_timeout": "900ms"
}
}
}'The ttft_timeout value is a duration string, such as "300ms", "1s", or "1.5s". The minimum is 300ms.
Auto model selection
Set model to auto and use extra_body.sort to tell the router how to rank models:
curl -X POST https://api.inworld.ai/v1/chat/completions \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "auto",
"messages": [{ "role": "user", "content": "Hello!" }],
"extra_body": {
"sort": ["price"]
}
}'This selects the cheapest available model. Available sort criteria: price, latency, throughput, intelligence, math, coding.
You can combine criteria. The router ranks models by the first criterion and uses the others to break ties:
curl -X POST https://api.inworld.ai/v1/chat/completions \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "auto",
"messages": [{ "role": "user", "content": "Hello!" }],
"extra_body": {
"sort": ["price", "latency"]
}
}'This picks the cheapest model, using latency as a tiebreaker.
Filtering models
Use extra_body.models to limit which models the router can choose. Use extra_body.ignore to exclude models or entire providers:
{
"model": "auto",
"messages": [{ "role": "user", "content": "Hello!" }],
"extra_body": {
"models": ["openai/gpt-5.2", "anthropic/claude-opus-4-6", "google-ai-studio/gemini-2.5-pro"],
"sort": ["latency"]
}
}{
"model": "auto",
"messages": [{ "role": "user", "content": "Hello!" }],
"extra_body": {
"ignore": ["openai", "anthropic/claude-opus-4-6"],
"sort": ["latency"]
}
}