Core Concepts
Responses API
Use the OpenAI Responses API format with any model on Inworld Router, with routing, fallbacks, and web search
Inworld Router serves the OpenAI Responses API at https://api.inworld.ai/v1/responses, alongside chat completions. Send the Responses format (input, instructions, typed output items, reasoning and tool items) to any model the router serves. OpenAI, Azure OpenAI, xAI and Fireworks models run on their providers' native Responses APIs. Other models, including Anthropic, Google and Inworld-hosted ones, are translated to their providers' APIs for you.
The router's model selection works here as it does on chat completions: provider/model ids, auto, saved routers (inworld/<router-id>), fallback models, and sort.
| Endpoint | Purpose |
|---|---|
POST /v1/responses | Create a response, streaming or not |
GET /v1/responses/{response_id} | Retrieve a stored response |
DELETE /v1/responses/{response_id} | Delete a stored response |
POST /v1/responses/{response_id}/cancel | Cancel a background response |
GET /v1/responses/{response_id}/input_items | List a stored response's input items |
The four {response_id} endpoints work only with responses stored by OpenAI models.
Quickstart
Run these examples on your server. Set INWORLD_API_KEY to the complete Base64 credentials copied from Portal or the CLI; see OpenAI compatibility.
import os
from openai import OpenAI
api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
base_url="https://api.inworld.ai/v1",
api_key=api_key,
default_headers={"Authorization": f"Basic {api_key}"},
)
response = client.responses.create(
model="anthropic/claude-sonnet-4-6",
instructions="You are terse.",
input="What is the capital of France?",
)
print(response.output_text)import OpenAI from 'openai';
const apiKey = process.env.INWORLD_API_KEY;
const client = new OpenAI({
baseURL: 'https://api.inworld.ai/v1',
apiKey,
defaultHeaders: { Authorization: `Basic ${apiKey}` },
});
const response = await client.responses.create({
model: 'anthropic/claude-sonnet-4-6',
instructions: 'You are terse.',
input: 'What is the capital of France?',
});
console.log(response.output_text);curl https://api.inworld.ai/v1/responses \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4-6",
"instructions": "You are terse.",
"input": "What is the capital of France?"
}'Request parameters
model and input are required. input is a string, which becomes a single user message, or an array of input items.
These Responses API parameters are passed to the model: instructions, tools, tool_choice, parallel_tool_calls, max_tool_calls, temperature, top_p, max_output_tokens, reasoning, text (including text.format for JSON schema output), stream, stream_options, store, background, previous_response_id, include, metadata, truncation, service_tier, prompt_cache_key, prompt_cache_retention, safety_identifier, user, and top_logprobs (0 to 20). Other top-level keys are forwarded to the provider unchanged, except the chat-only and unsupported parameters below and the routing options.
When the selected model does not support temperature, top_p, or reasoning.effort, the router drops that parameter rather than failing the request, and lists it in routing.attempts[].warnings (reasoning.effort appears there as reasoning_effort). top_logprobs is not supported and is dropped.
Not supported
| Parameter | Result |
|---|---|
conversation | 400 with code: "unsupported_parameter". Continue a conversation with previous_response_id, or send the history in input |
prompt (hosted prompt templates) | 400 with code: "unsupported_parameter". Send instructions and input directly |
item_reference input items | 400 with code: "unsupported_parameter" |
Chat Completions parameters: messages, max_tokens, max_completion_tokens, response_format, n, logprobs, stop, seed, frequency_penalty, presence_penalty, logit_bias, reasoning_effort, web_search_options, modalities, audio, prompt_variables, thinking, and other chat-only fields | 400 naming the parameter. Use the Responses equivalent (for example max_output_tokens, text.format, reasoning.effort), or send the request to /v1/chat/completions |
A Content-Type other than application/json is refused with 415.
Routing options
These Inworld options go at the top level of the body or inside extra_body (the OpenAI SDK's field for extra parameters). When both set one, the top level wins.
| Option | Description |
|---|---|
models | Fallback models, tried in order when the primary fails. See specific models |
ignore | Models to exclude from auto selection |
sort | Ranking for auto and fallbacks, for example ["price"] or ["latency"]. See auto model selection |
provider | {"order": [...], "allow_fallbacks": true} to control which providers serve the model |
fallback | {"ttft_timeout": "2s"} to fall back when the first token is slower than this duration (at least 300ms) |
session | A session id (up to 128 printable ASCII characters, not default) for grouping the requests of one conversation |
web_search | Router-run web search; see below |
Unrecognized keys inside extra_body are refused with 400. auto cannot appear in models, ignore, or provider.order.
With a saved router (model: "inworld/<router-id>"), the router's model selection replaces models, ignore, sort and provider, and its generation settings apply only when the request sets none of temperature, top_p, max_output_tokens, top_logprobs and reasoning.effort. If the request sets any of them, none of the router's generation settings apply. Router settings that have no Responses equivalent (audio output, image settings, prompt compression) are skipped with a warning in routing.attempts[].warnings.
Some Inworld-hosted models (inworld/models/...) are available only on /v1/chat/completions. For those, the request is refused with 400 and a message pointing there.
Response
The response is an OpenAI Response object, with these differences:
routingis added to every response created withPOST /v1/responses. It has the same fields asmetadataon chat completions:attempts(each with the provider-qualifiedmodel,success,status_code,duration_ms,time_to_first_token_msandwarnings),generation_id, andtotal_duration_ms, plusroute_idandvariant_idwhen a saved router chose the model. Themetadatafield is left for your own request metadata, as in OpenAI's API.modelis the model id the provider reports, for examplegpt-5.4. For the provider-qualified id (openai/gpt-5.4), readrouting.attempts[].model.- Retrieve, delete, cancel and input-item responses carry no
routingobject.
A response from openai/gpt-5.4:
{
"id": "resp_1",
"object": "response",
"model": "gpt-5.4",
"status": "completed",
"output": [
{
"id": "msg_1",
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Paris." }]
}
],
"usage": { "input_tokens": 21, "output_tokens": 3, "total_tokens": 24 },
"routing": {
"attempts": [
{ "model": "openai/gpt-5.4", "success": true, "time_to_first_token_ms": 80 }
],
"generation_id": "5f0c1c3e-3b53-4d8e-9f53-0c2b8a4de6b1",
"total_duration_ms": 412
}
}Streaming
With stream: true, the response is a server-sent event stream of named events (response.created, response.output_text.delta, response.completed, and so on), as in OpenAI's API. The OpenAI SDKs read it unchanged.
- The stream ends after the terminal event (
response.completed,response.incomplete, orresponse.failed). There is nodata: [DONE]line. routingis included on events that carry theresponseobject, such asresponse.createdandresponse.completed.- A failure after the stream starts is sent as an
errorevent,{"type": "error", "code": ..., "message": ..., "param": ..., "sequence_number": ...}, and the stream ends.
Web search
Two ways to ground a response with web results:
- Router-run search. Set
web_search, as on chat completions:{"engine": "exa", "max_results": 3, "max_steps": 1}. The router gives the model a search tool, runs the searches, and returns one response withurl_citationannotations on the final message. Usage is summed across the search steps. Router-run search is refused withbackgroundorprevious_response_id. - Native search. Set
web_search: {"engine": "native"}, or pass the hosted tool directly:"tools": [{"type": "web_search"}]. The provider runs the search. Native search works with background mode and continuation.
Stored responses and continuation
| Provider | store: true and continuation | background: true | Retrieve, delete, cancel, input items |
|---|---|---|---|
| OpenAI | Yes | Yes | Yes |
| Azure OpenAI | Yes | No (400) | No (400, code: "unsupported_operation") |
| Every other provider | No (400) | No (400) | No |
Store is off unless you ask for it. With Inworld-managed credentials, a response is stored only when the request sets store: true or background: true. That differs from OpenAI's API, where store defaults to true. With your own OpenAI or Azure OpenAI key, the provider's default applies. To continue a conversation with previous_response_id, set store: true on every turn you will continue from.
first = client.responses.create(
model="openai/gpt-5.4",
input="Pick a number between 1 and 10.",
store=True,
)
second = client.responses.create(
model="openai/gpt-5.4",
input="Double it.",
previous_response_id=first.id,
store=True,
)A continuation (previous_response_id) must name one model on the provider that stored the response, OpenAI or Azure OpenAI. Sending fallback models with previous_response_id returns 400. A saved router can continue only a response that this endpoint stored for your workspace.
With Inworld-managed credentials, a stored response id is usable only by the workspace that created it, for 7 days. An unknown id, or one created by another workspace, returns 404. With your own provider key, the provider's own storage rules apply.
Background mode
background: true runs the response asynchronously on OpenAI models. Poll it with retrieve, or stop it with cancel. Background responses are billed for the tokens they use, including a response you cancel.
With Inworld-managed credentials, background mode must be enabled for your workspace. Otherwise the request is refused with 400. Contact support to enable it. With your own OpenAI key, background mode is always available.
Zero data retention
For a workspace with zero data retention, nothing is stored. store: true, background and previous_response_id are refused with 400 (code: "unsupported_parameter"), and so are the four {response_id} endpoints.
Errors
Errors use the router's format: {"error": {"message", "type", "code", "param"}}, with metadata.attempts when models were tried.
| Status | Common causes |
|---|---|
400 | Invalid JSON, missing model or input, an unsupported parameter, store/background/previous_response_id on a model that cannot store |
401 | Missing or invalid API key. The body is plain text (Unauthorized) |
402 | code: "insufficient_credits" |
404 | Unknown saved router, no model matches the request, or an unknown response id |
415 | Content-Type is not application/json |
429 | Rate limited. Retry after Retry-After |
503 | Every candidate model is unavailable, or the service is temporarily unavailable. Retry after Retry-After |
504 | The request ran out of time. Retry after Retry-After |
When every candidate model fails, the status is the last failure's, and the message lists each model's error.