Responses
Create response
Create a model response in the OpenAI Responses API format
/v1/responsesSend the OpenAI Responses API format to any model Inworld Router serves: provider/model ids, auto, or a saved router (inworld/<router-id>), with fallback models, sort, and web search. The OpenAI SDKs work with the base URL set to https://api.inworld.ai/v1 and a Basic Authorization header.
Every response carries a routing object with the models tried. Stored responses and previous_response_id work on OpenAI and Azure OpenAI models only, and store defaults to false. See the Responses API guide for supported parameters, routing options, streaming, and stored responses.
Authorizationstringrequired
Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.
Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.
modelstringrequired
A provider/model id (for example openai/gpt-5.4), auto for automatic selection, a saved router (inworld/<router-id>), or an Inworld-hosted model (inworld/models/<name>).
inputoneOfrequired
A string (sent as one user message) or a non-empty array of input items (messages, function call outputs, reasoning items). item_reference items are not supported.
instructionsstring
A system or developer message placed before the input.
toolsobject[]
Tools the model may call: function tools, and hosted tools such as {"type": "web_search"} where the provider supports them.
tool_choiceoneOf
How the model chooses tools: auto, none, required, or a specific tool.
parallel_tool_callsboolean
max_tool_callsinteger
temperaturenumber
Dropped, with a warning in routing.attempts[].warnings, for models that do not support it.
top_pnumber
Dropped, with a warning, for models that do not support it.
max_output_tokensinteger
Upper bound on output tokens, including reasoning tokens.
reasoningobject
Reasoning configuration. effort is dropped, with a warning, for models that do not support it.
Show child attributes
effortenum<string>
Available options:noneminimallowmediumhighxhigh
summarystring
textobject
Text output configuration, including format ({"type": "json_schema", ...} for structured output) and verbosity.
streambooleandefault: false
stream_optionsobject
storeboolean
Store the response so it can be continued (OpenAI and Azure OpenAI) or retrieved (OpenAI). With Inworld-managed credentials, defaults to false, unlike OpenAI's API; with your own provider key, the provider's default applies. background: true also stores the response. Refused under zero data retention.
backgroundboolean
Run the response asynchronously. OpenAI models only. With Inworld-managed credentials, background mode must be enabled for your workspace; contact support. Refused under zero data retention.
previous_response_idstring
Continue from a response this workspace stored with store: true. Name one model on the provider that stored it (OpenAI or Azure OpenAI), without fallback models. Refused under zero data retention.
includestring[]
metadataobject
Your key-value metadata, as in OpenAI's API.
truncationenum<string>
Available options:autodisabled
top_logprobsinteger
Not supported. The parameter is dropped with a warning.
service_tierstring
prompt_cache_keystring
prompt_cache_retentionstring
safety_identifierstring
userstring
modelsstring[]
Inworld: fallback models, tried in order when the primary fails. auto is not allowed here.
ignorestring[]
Inworld: models to exclude from selection.
sortstring[]
Inworld: ranking criteria for auto and fallbacks, for example ["price"] or ["latency"].
providerobject
Inworld: provider preferences.
Show child attributes
orderstring[]
allow_fallbacksboolean
fallbackobject
Inworld: {"ttft_timeout": ...} falls back to the next model when the first token is slow.
Show child attributes
ttft_timeoutstring
sessionstring
Inworld: an optional session id for grouping the requests of one conversation. Printable ASCII; default is reserved.
web_searchobject
Web search configuration, at the top level or in extra_body. With a router-run engine, the model calls a search tool in a loop and the router returns a grounded answer; native uses the provider's built-in search. See Web search.
Show child attributes
engineenum<string>default: "exa"
Search backend: exa, brave or google, or native for the provider's built-in search.
Available options:exabravegooglenative
max_resultsintegerdefault: 3
Search results per search call, 1 to 10.
max_stepsintegerdefault: 1
Maximum search rounds before the final answer, 1 to 5.
extra_bodyobject
Inworld routing options (models, ignore, sort, provider, fallback, session, web_search) for SDKs that send extra parameters here. Unrecognized keys are refused with 400.
idstring
objectstring
created_atinteger
statusenum<string>
Available options:completedfailedin_progressqueuedcancelledincomplete
modelstring
The model id the provider reports, for example gpt-5.4. The provider-qualified id is in routing.attempts[].model.
outputobject[]
Output items: message, reasoning, function_call, and hosted tool items.
usageobject
Show child attributes
input_tokensinteger
output_tokensinteger
total_tokensinteger
input_tokens_detailsobject
output_tokens_detailsobject
errorobject
incomplete_detailsobject
routingobject
Inworld routing information. Present on responses from POST /v1/responses, not on retrieved responses.
Show child attributes
attemptsobject[]
Show child attributes
modelstring
The model identifier that was attempted.
successboolean
Whether this attempt succeeded.
time_to_first_token_msinteger
Time to receive the first token in milliseconds.
status_codeinteger
HTTP status of a failed attempt.
errorstring
Why the attempt failed.
duration_msnumber
Duration of the attempt in milliseconds.
warningsstring[]
Notes about the attempt, for example parameters skipped because the model does not support them.
generation_idstring
route_idstring
The saved router's route that was chosen, when model is a saved router.
variant_idstring
The saved router's variant that was chosen, when model is a saved router.
total_duration_msinteger