Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Responses

Create response

Create a model response in the OpenAI Responses API format

POST/v1/responses

Send the OpenAI Responses API format to any model Inworld Router serves: provider/model ids, auto, or a saved router (inworld/<router-id>), with fallback models, sort, and web search. The OpenAI SDKs work with the base URL set to https://api.inworld.ai/v1 and a Basic Authorization header.

Every response carries a routing object with the models tried. Stored responses and previous_response_id work on OpenAI and Azure OpenAI models only, and store defaults to false. See the Responses API guide for supported parameters, routing options, streaming, and stored responses.

Authorizations

Authorizationstringrequired

Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.

Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Body

application/json

modelstringrequired

A provider/model id (for example openai/gpt-5.4), auto for automatic selection, a saved router (inworld/<router-id>), or an Inworld-hosted model (inworld/models/<name>).

inputoneOfrequired

A string (sent as one user message) or a non-empty array of input items (messages, function call outputs, reasoning items). item_reference items are not supported.

instructionsstring

A system or developer message placed before the input.

toolsobject[]

Tools the model may call: function tools, and hosted tools such as {"type": "web_search"} where the provider supports them.

tool_choiceoneOf

How the model chooses tools: auto, none, required, or a specific tool.

parallel_tool_callsboolean

max_tool_callsinteger

temperaturenumber

Dropped, with a warning in routing.attempts[].warnings, for models that do not support it.

top_pnumber

Dropped, with a warning, for models that do not support it.

max_output_tokensinteger

Upper bound on output tokens, including reasoning tokens.

reasoningobject

Reasoning configuration. effort is dropped, with a warning, for models that do not support it.

Show child attributes

effortenum<string>

Available options:noneminimallowmediumhighxhigh

summarystring

textobject

Text output configuration, including format ({"type": "json_schema", ...} for structured output) and verbosity.

streambooleandefault: false

stream_optionsobject

storeboolean

Store the response so it can be continued (OpenAI and Azure OpenAI) or retrieved (OpenAI). With Inworld-managed credentials, defaults to false, unlike OpenAI's API; with your own provider key, the provider's default applies. background: true also stores the response. Refused under zero data retention.

backgroundboolean

Run the response asynchronously. OpenAI models only. With Inworld-managed credentials, background mode must be enabled for your workspace; contact support. Refused under zero data retention.

previous_response_idstring

Continue from a response this workspace stored with store: true. Name one model on the provider that stored it (OpenAI or Azure OpenAI), without fallback models. Refused under zero data retention.

includestring[]

metadataobject

Your key-value metadata, as in OpenAI's API.

truncationenum<string>

Available options:autodisabled

top_logprobsinteger

Not supported. The parameter is dropped with a warning.

service_tierstring

prompt_cache_keystring

prompt_cache_retentionstring

safety_identifierstring

userstring

modelsstring[]

Inworld: fallback models, tried in order when the primary fails. auto is not allowed here.

ignorestring[]

Inworld: models to exclude from selection.

sortstring[]

Inworld: ranking criteria for auto and fallbacks, for example ["price"] or ["latency"].

providerobject

Inworld: provider preferences.

Show child attributes

orderstring[]

allow_fallbacksboolean

fallbackobject

Inworld: {"ttft_timeout": ...} falls back to the next model when the first token is slow.

Show child attributes

ttft_timeoutstring

sessionstring

Inworld: an optional session id for grouping the requests of one conversation. Printable ASCII; default is reserved.

web_searchobject

Web search configuration, at the top level or in extra_body. With a router-run engine, the model calls a search tool in a loop and the router returns a grounded answer; native uses the provider's built-in search. See Web search.

Show child attributes

engineenum<string>default: "exa"

Search backend: exa, brave or google, or native for the provider's built-in search.

Available options:exabravegooglenative

max_resultsintegerdefault: 3

Search results per search call, 1 to 10.

max_stepsintegerdefault: 1

Maximum search rounds before the final answer, 1 to 5.

extra_bodyobject

Inworld routing options (models, ignore, sort, provider, fallback, session, web_search) for SDKs that send extra parameters here. Unrecognized keys are refused with 400.

Response

200 - application/json

idstring

objectstring

created_atinteger

statusenum<string>

Available options:completedfailedin_progressqueuedcancelledincomplete

modelstring

The model id the provider reports, for example gpt-5.4. The provider-qualified id is in routing.attempts[].model.

outputobject[]

Output items: message, reasoning, function_call, and hosted tool items.

usageobject

Show child attributes

input_tokensinteger

output_tokensinteger

total_tokensinteger

input_tokens_detailsobject

output_tokens_detailsobject

errorobject

incomplete_detailsobject

routingobject

Inworld routing information. Present on responses from POST /v1/responses, not on retrieved responses.

Show child attributes

attemptsobject[]

Show child attributes

modelstring

The model identifier that was attempted.

successboolean

Whether this attempt succeeded.

time_to_first_token_msinteger

Time to receive the first token in milliseconds.

status_codeinteger

HTTP status of a failed attempt.

errorstring

Why the attempt failed.

duration_msnumber

Duration of the attempt in milliseconds.

warningsstring[]

Notes about the attempt, for example parameters skipped because the model does not support them.

generation_idstring

route_idstring

The saved router's route that was chosen, when model is a saved router.

variant_idstring

The saved router's variant that was chosen, when model is a saved router.

total_duration_msinteger