Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Routing

Request-level routing

Route requests without creating a router

With the Chat Completions API, you can route requests without creating a router. In a single request, you can call a specific model, add fallbacks, or let the router choose a model for you.

Use request-level routing to:

  • Call a specific model through a unified API without setting up a router
  • Prototype or benchmark before choosing a router configuration

For conditional routing, A/B testing with weighted variants, or shared prompt templates, set up a router.

Authentication

Run these examples on your server. Set INWORLD_API_KEY to the Base64 credentials copied from Portal or the CLI. Use the value as-is; don't Base64-encode it again. A Standard key can call chat completions. You only need Router Write permission to change router configuration.

The direct HTTP examples use Basic. Keep API keys on your server. Browser and mobile clients should use a token created by your backend or a server proxy.

Direct model call

Set model to a provider/model identifier:

bash
curl -X POST https://api.inworld.ai/v1/chat/completions \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "openai/gpt-5.2",
    "messages": [{ "role": "user", "content": "Hello!" }]
  }'

This calls the specified model through LLM Router's unified API, with no routing logic.

Fallbacks

Add fallback models via extra_body.models. If the primary model fails, the router automatically tries the next model in the list:

bash
curl -X POST https://api.inworld.ai/v1/chat/completions \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "openai/gpt-5.2",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "extra_body": {
      "models": ["anthropic/claude-opus-4-6", "google-ai-studio/gemini-2.5-pro"]
    }
  }'

In this example, the router tries gpt-5.2 first, then Claude Opus, then Gemini Pro. The response's metadata.attempts array shows which models it tried.

Fallback by first token timeout

Set a time to first token (TTFT) timeout to trigger fallback when a model is slow. If the model doesn't return its first token in time, the router cancels the request and tries the next model.

Use this when your application needs a fast response and can try another model instead of waiting.

Set extra_body.fallback.ttft_timeout in your request:

bash
curl -X POST https://api.inworld.ai/v1/chat/completions \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "openai/gpt-5.2",
    "messages": [{ "role": "user", "content": "Hello" }],
    "extra_body": {
      "models": ["openai/gpt-4o", "google-ai-studio/gemini-2.5-pro"],
      "fallback": {
        "ttft_timeout": "900ms"
      }
    }
  }'

The ttft_timeout value is a duration string, such as "300ms", "1s", or "1.5s". The minimum is 300ms.

Auto model selection

Set model to auto and use extra_body.sort to tell the router how to rank models:

bash
curl -X POST https://api.inworld.ai/v1/chat/completions \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "auto",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "extra_body": {
      "sort": ["price"]
    }
  }'

This selects the cheapest available model. Available sort criteria: price, latency, throughput, intelligence, math, coding.

You can combine criteria. The router ranks models by the first criterion and uses the others to break ties:

bash
curl -X POST https://api.inworld.ai/v1/chat/completions \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "auto",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "extra_body": {
      "sort": ["price", "latency"]
    }
  }'

This picks the cheapest model, using latency as a tiebreaker.

Filtering models

Use extra_body.models to limit which models the router can choose. Use extra_body.ignore to exclude models or entire providers:

Restrict to specific models
{
  "model": "auto",
  "messages": [{ "role": "user", "content": "Hello!" }],
  "extra_body": {
    "models": ["openai/gpt-5.2", "anthropic/claude-opus-4-6", "google-ai-studio/gemini-2.5-pro"],
    "sort": ["latency"]
  }
}