Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

API formats

Anthropic compatibility

LLM Router is available through Anthropic-compatible API endpoints, so you can use the Anthropic SDK and tools like Claude Code, while benefiting from LLM Router's multi-provider routing, fallbacks, and cost optimization.

Credential handling

Run these examples on your server. Set INWORLD_API_KEY to the complete Base64 credentials copied from Portal or the CLI, without encoding them again. The examples set an explicit Authorization: Basic ... header. Browser/mobile clients need a backend-minted token, never the server API key.

Endpoints

The Anthropic-compatible endpoints are:

EndpointPurpose
POST https://api.inworld.ai/v1/messagesCreate a message, streaming or not
POST https://api.inworld.ai/v1/messages/count_tokensCount input tokens

When using the Anthropic SDK, set the base URL to https://api.inworld.ai. The SDK automatically appends the /v1/messages paths.

Authentication

For a direct HTTP request, use the copied Base64 credential with Basic:

Authorization: Basic <base64-credential>

In the SDK examples below, auth_token / authToken gives the SDK the credential it requires, and default_headers / defaultHeaders sends that credential as the Basic authorization header, in place of the Bearer header the SDK would otherwise send. Use both settings as shown.

The x-api-key header is not read, so an Anthropic-style api_key / ANTHROPIC_API_KEY setup fails authentication. The examples set api_key / apiKey to None / null, and the Python example also omits the X-Api-Key header, so an ANTHROPIC_API_KEY in your environment is not sent.

Anthropic SDK

Below is an example request using Anthropic's SDK

Python
import os
import anthropic

api_key = os.environ["INWORLD_API_KEY"]

client = anthropic.Anthropic(
    base_url="https://api.inworld.ai",
    api_key=None,
    auth_token=api_key,
    default_headers={"Authorization": f"Basic {api_key}", "X-Api-Key": anthropic.omit},
)

message = client.messages.create(
    model="inworld/<router-id>",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Explain how neural networks learn."}
    ]
)

print(message.content[0].text)

Response

The response follows the Anthropic Messages API format:

json
{
  "content": [
    {"text": "Neural networks learn through...", "type": "text"}
  ],
  "id": "msg-...",
  "model": "anthropic/claude-opus-4-6",
  "role": "assistant",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "type": "message",
  "usage": {
    "input_tokens": 12,
    "output_tokens": 150,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0
  },
  "metadata": {
    "attempts": [
      {"model": "anthropic/claude-opus-4-6", "success": true, "time_to_first_token_ms": 1152}
    ],
    "generation_id": "019b...",
    "total_duration_ms": 1480,
    "reasoning": "Using specified model: 'anthropic/claude-opus-4-6' - success"
  }
}

On non-streaming responses, LLM Router adds a metadata field containing routing information โ€” attempt history, timing, and, when applicable, route_id, variant_id, and reasoning. This field is not part of the standard Anthropic response format, but it does not break Anthropic SDK parsing.

Supported and unsupported features

/v1/messages translates each request into the router's chat format, routes it, and translates the result back. Fields and content blocks that have no translation are dropped silently, without an error.

Request parameters

SupportedNotes
model, messages, max_tokens, streammodel accepts provider/model, a bare Claude name, auto, or inworld/<router-id>
systemString or array of text blocks; multiple blocks are joined with a space unless a block carries cache_control
temperature, top_p, stop_sequences
tools, tool_choiceTools use name, description, and input_schema. tool_choice supports auto, any, and tool; none is ignored, so the tools stay available
metadatametadata.user_id is used as the end-user ID
cache_controlTop-level, and on system, message, content, and tool blocks
thinking{"type": "enabled", "budget_tokens": N} sets the reasoning token budget

The router extensions models, sort, ignore, reasoning, prompt_variables, and session are accepted as top-level fields. To send other router extensions, such as provider or web_search_options, nest them in an extra_body object in the JSON body. The Anthropic SDK's extra_body option adds its keys at the top level, so pass the top-level extensions directly and nest the others one level deeper, as in extra_body={"extra_body": {"provider": {...}}}:

python
message = client.messages.create(
    model="anthropic/claude-opus-4-6",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello!"}],
    extra_body={
        "models": ["openai/gpt-5.4", "google-ai-studio/gemini-2.5-flash"],
        "sort": ["price"],
    },
)

Refused: audio, top level or in extra_body, returns 400 with code: "unsupported_parameter", because an Anthropic response has no audio block. For voice responses, use LLM + TTS on /v1/chat/completions.

Dropped: top_k, service_tier, mcp_servers, container, and any other top-level field not listed above. tool_choice.disable_parallel_tool_use is also dropped. Anthropic server tools and special tool types (web search, code execution, computer use, and so on) are forwarded as ordinary function definitions built from name, description, and input_schema; the router does not execute them.

Content blocks

Block typeHandling
textConverted. Multiple text blocks in one message are joined with a space, unless a block carries cache_control or the message contains an image; then they stay separate content parts
imageConverted, from base64 or url sources
tool_useConverted to a tool call. Repeated IDs within a message are de-duplicated
tool_resultConverted to a tool message. Only a string content or the first text block is kept; images inside a tool result are dropped
document (including PDFs), search_result, thinking, redacted_thinking, server tool result blocksDropped

Because thinking blocks in conversation history are dropped, multi-turn reasoning with tool calls does not round-trip thinking signatures.

Response differences

  • stop_sequence is always null. When a stop sequence ends generation, stop_reason is end_turn.
  • Non-streaming stop_reason is end_turn, tool_use, or max_tokens. Streaming responses report only end_turn or tool_use.
  • Reasoning output is returned as a thinking block without a signature.
  • A non-streaming response always contains a text block, which is empty when the model only calls tools. A streaming response includes a text block only when the model returns text.
  • usage always includes cache_read_input_tokens and cache_creation_input_tokens, which are 0 when nothing was cached.

Errors

Errors that the router raises before a response starts use the OpenAI-style error body with the matching HTTP status, not Anthropic's {"type": "error", ...} envelope:

json
{
  "error": {
    "message": "'model' field is required",
    "type": "invalid_request_error"
  }
}

The Anthropic SDKs still raise the usual status-based exceptions, but code that parses the error body directly should read error.message and error.type. Errors during a stream are sent as an Anthropic error event with type api_error.

Authentication and permission errors (401, 403), and a 404 for an unknown path under /v1/messages, have a plain-text body such as Unauthorized instead.

Count tokens

POST /v1/messages/count_tokens returns the input token count for a request without generating a response:

bash
curl -X POST https://api.inworld.ai/v1/messages/count_tokens \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-opus-4-6",
    "messages": [
      {"role": "user", "content": "Explain how neural networks learn."}
    ]
  }'
json
{"input_tokens": 14}

The count is computed for the named model only, without routing or fallbacks. Keep these limits in mind:

  • model is required and must be a specific model. A bare name is treated as an Anthropic model (claude-opus-4-6 becomes anthropic/claude-opus-4-6). auto and inworld/<router-id> are not supported.
  • Only model, messages, tools, and tool_choice are read. The top-level system prompt is not counted; to include it, add it as a message with role: "system".
  • Only text is counted. For array content, the text of each block is counted; images, tool_use, and tool_result blocks are not.
  • Tools must use the OpenAI shape ({"type": "function", "function": {"name", "description", "parameters"}}) to be counted. Anthropic-shaped tools (name, input_schema) are not counted accurately.
  • Errors: a missing model or invalid JSON returns 400, and a Content-Type other than application/json returns 415. Any failure while counting, including an unknown model, returns 500 with type api_error.

Next Steps