Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Core Concepts

Responses API

Use the OpenAI Responses API format with any model on Inworld Router, with routing, fallbacks, and web search

Inworld Router serves the OpenAI Responses API at https://api.inworld.ai/v1/responses, alongside chat completions. Send the Responses format (input, instructions, typed output items, reasoning and tool items) to any model the router serves. OpenAI, Azure OpenAI, xAI and Fireworks models run on their providers' native Responses APIs. Other models, including Anthropic, Google and Inworld-hosted ones, are translated to their providers' APIs for you.

The router's model selection works here as it does on chat completions: provider/model ids, auto, saved routers (inworld/<router-id>), fallback models, and sort.

EndpointPurpose
POST /v1/responsesCreate a response, streaming or not
GET /v1/responses/{response_id}Retrieve a stored response
DELETE /v1/responses/{response_id}Delete a stored response
POST /v1/responses/{response_id}/cancelCancel a background response
GET /v1/responses/{response_id}/input_itemsList a stored response's input items

The four {response_id} endpoints work only with responses stored by OpenAI models.

Quickstart

Run these examples on your server. Set INWORLD_API_KEY to the complete Base64 credentials copied from Portal or the CLI; see OpenAI compatibility.

Python
import os
from openai import OpenAI

api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
    base_url="https://api.inworld.ai/v1",
    api_key=api_key,
    default_headers={"Authorization": f"Basic {api_key}"},
)

response = client.responses.create(
    model="anthropic/claude-sonnet-4-6",
    instructions="You are terse.",
    input="What is the capital of France?",
)
print(response.output_text)

Request parameters

model and input are required. input is a string, which becomes a single user message, or an array of input items.

These Responses API parameters are passed to the model: instructions, tools, tool_choice, parallel_tool_calls, max_tool_calls, temperature, top_p, max_output_tokens, reasoning, text (including text.format for JSON schema output), stream, stream_options, store, background, previous_response_id, include, metadata, truncation, service_tier, prompt_cache_key, prompt_cache_retention, safety_identifier, user, and top_logprobs (0 to 20). Other top-level keys are forwarded to the provider unchanged, except the chat-only and unsupported parameters below and the routing options.

When the selected model does not support temperature, top_p, or reasoning.effort, the router drops that parameter rather than failing the request, and lists it in routing.attempts[].warnings (reasoning.effort appears there as reasoning_effort). top_logprobs is not supported and is dropped.

Not supported

ParameterResult
conversation400 with code: "unsupported_parameter". Continue a conversation with previous_response_id, or send the history in input
prompt (hosted prompt templates)400 with code: "unsupported_parameter". Send instructions and input directly
item_reference input items400 with code: "unsupported_parameter"
Chat Completions parameters: messages, max_tokens, max_completion_tokens, response_format, n, logprobs, stop, seed, frequency_penalty, presence_penalty, logit_bias, reasoning_effort, web_search_options, modalities, audio, prompt_variables, thinking, and other chat-only fields400 naming the parameter. Use the Responses equivalent (for example max_output_tokens, text.format, reasoning.effort), or send the request to /v1/chat/completions

A Content-Type other than application/json is refused with 415.

Routing options

These Inworld options go at the top level of the body or inside extra_body (the OpenAI SDK's field for extra parameters). When both set one, the top level wins.

OptionDescription
modelsFallback models, tried in order when the primary fails. See specific models
ignoreModels to exclude from auto selection
sortRanking for auto and fallbacks, for example ["price"] or ["latency"]. See auto model selection
provider{"order": [...], "allow_fallbacks": true} to control which providers serve the model
fallback{"ttft_timeout": "2s"} to fall back when the first token is slower than this duration (at least 300ms)
sessionA session id (up to 128 printable ASCII characters, not default) for grouping the requests of one conversation
web_searchRouter-run web search; see below

Unrecognized keys inside extra_body are refused with 400. auto cannot appear in models, ignore, or provider.order.

With a saved router (model: "inworld/<router-id>"), the router's model selection replaces models, ignore, sort and provider, and its generation settings apply only when the request sets none of temperature, top_p, max_output_tokens, top_logprobs and reasoning.effort. If the request sets any of them, none of the router's generation settings apply. Router settings that have no Responses equivalent (audio output, image settings, prompt compression) are skipped with a warning in routing.attempts[].warnings.

Some Inworld-hosted models (inworld/models/...) are available only on /v1/chat/completions. For those, the request is refused with 400 and a message pointing there.

Response

The response is an OpenAI Response object, with these differences:

  • routing is added to every response created with POST /v1/responses. It has the same fields as metadata on chat completions: attempts (each with the provider-qualified model, success, status_code, duration_ms, time_to_first_token_ms and warnings), generation_id, and total_duration_ms, plus route_id and variant_id when a saved router chose the model. The metadata field is left for your own request metadata, as in OpenAI's API.
  • model is the model id the provider reports, for example gpt-5.4. For the provider-qualified id (openai/gpt-5.4), read routing.attempts[].model.
  • Retrieve, delete, cancel and input-item responses carry no routing object.

A response from openai/gpt-5.4:

json
{
  "id": "resp_1",
  "object": "response",
  "model": "gpt-5.4",
  "status": "completed",
  "output": [
    {
      "id": "msg_1",
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Paris." }]
    }
  ],
  "usage": { "input_tokens": 21, "output_tokens": 3, "total_tokens": 24 },
  "routing": {
    "attempts": [
      { "model": "openai/gpt-5.4", "success": true, "time_to_first_token_ms": 80 }
    ],
    "generation_id": "5f0c1c3e-3b53-4d8e-9f53-0c2b8a4de6b1",
    "total_duration_ms": 412
  }
}

Streaming

With stream: true, the response is a server-sent event stream of named events (response.created, response.output_text.delta, response.completed, and so on), as in OpenAI's API. The OpenAI SDKs read it unchanged.

  • The stream ends after the terminal event (response.completed, response.incomplete, or response.failed). There is no data: [DONE] line.
  • routing is included on events that carry the response object, such as response.created and response.completed.
  • A failure after the stream starts is sent as an error event, {"type": "error", "code": ..., "message": ..., "param": ..., "sequence_number": ...}, and the stream ends.

Two ways to ground a response with web results:

  • Router-run search. Set web_search, as on chat completions: {"engine": "exa", "max_results": 3, "max_steps": 1}. The router gives the model a search tool, runs the searches, and returns one response with url_citation annotations on the final message. Usage is summed across the search steps. Router-run search is refused with background or previous_response_id.
  • Native search. Set web_search: {"engine": "native"}, or pass the hosted tool directly: "tools": [{"type": "web_search"}]. The provider runs the search. Native search works with background mode and continuation.

Stored responses and continuation

Providerstore: true and continuationbackground: trueRetrieve, delete, cancel, input items
OpenAIYesYesYes
Azure OpenAIYesNo (400)No (400, code: "unsupported_operation")
Every other providerNo (400)No (400)No

Store is off unless you ask for it. With Inworld-managed credentials, a response is stored only when the request sets store: true or background: true. That differs from OpenAI's API, where store defaults to true. With your own OpenAI or Azure OpenAI key, the provider's default applies. To continue a conversation with previous_response_id, set store: true on every turn you will continue from.

python
first = client.responses.create(
    model="openai/gpt-5.4",
    input="Pick a number between 1 and 10.",
    store=True,
)
second = client.responses.create(
    model="openai/gpt-5.4",
    input="Double it.",
    previous_response_id=first.id,
    store=True,
)

A continuation (previous_response_id) must name one model on the provider that stored the response, OpenAI or Azure OpenAI. Sending fallback models with previous_response_id returns 400. A saved router can continue only a response that this endpoint stored for your workspace.

With Inworld-managed credentials, a stored response id is usable only by the workspace that created it, for 7 days. An unknown id, or one created by another workspace, returns 404. With your own provider key, the provider's own storage rules apply.

Background mode

background: true runs the response asynchronously on OpenAI models. Poll it with retrieve, or stop it with cancel. Background responses are billed for the tokens they use, including a response you cancel.

With Inworld-managed credentials, background mode must be enabled for your workspace. Otherwise the request is refused with 400. Contact support to enable it. With your own OpenAI key, background mode is always available.

Zero data retention

For a workspace with zero data retention, nothing is stored. store: true, background and previous_response_id are refused with 400 (code: "unsupported_parameter"), and so are the four {response_id} endpoints.

Errors

Errors use the router's format: {"error": {"message", "type", "code", "param"}}, with metadata.attempts when models were tried.

StatusCommon causes
400Invalid JSON, missing model or input, an unsupported parameter, store/background/previous_response_id on a model that cannot store
401Missing or invalid API key. The body is plain text (Unauthorized)
402code: "insufficient_credits"
404Unknown saved router, no model matches the request, or an unknown response id
415Content-Type is not application/json
429Rate limited. Retry after Retry-After
503Every candidate model is unavailable, or the service is temporarily unavailable. Retry after Retry-After
504The request ran out of time. Retry after Retry-After

When every candidate model fails, the status is the last failure's, and the message lists each model's error.