Chat Completions
Create chat completion
Generate a response for the given chat conversation
/v1/chat/completionsCall 100+ models from various providers directly through our unified API, or set model to auto for automatic model selection based on criteria like price, latency, or performance.
For more advanced routing — such as conditional routing, A/B testing across variants, and reusable configurations — create a router and reference it via the model field (e.g., inworld/my-router).
For web-grounded answers, use extra_body.web_search.
For voice output, see Voice responses, including optional word or character timestamps.
Authorizationstringrequired
Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.
Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.
modelstringrequired
The model to use, which can be:
- A model id (e.g.,
deepseek-v4-flash). The best provider is automatically selected by latency, or you can control provider selection viaextra_body.provider. See Models for available models. - A provider-prefixed model id (e.g.,
openai/gpt-5.4). This specifies the provider and model to use. autofor automatic model selection based on criteria like price, latency, or intelligence- A router, which is specified by
inworld/<router-name>. The routernamemust be prefixed byinworld/.
messagesobject[]required
A list of messages comprising the conversation so far.
If using a router where a prompt is specified, these messages will be appended to the prompt.
Show child attributes
roleenum<string>required
The role of the message author.
Available options:systemuserassistanttool
contentoneOf
The content of the message. Can be a string for text-only messages, an array of content parts for multimodal messages, or null for assistant messages with tool_calls.
tool_callsobject[]
Tool calls generated by the model (assistant messages only).
Show child attributes
idstringrequired
ID of the tool call.
typeenum<string>required
The type of the tool call. Always 'function'.
Available options:function
functionobjectrequired
The function that the model called.
Show child attributes
namestringrequired
The name of the function to call.
argumentsstringrequired
The arguments to call the function with, as a JSON string.
tool_call_idstring
Tool call ID this message is responding to (tool role only).
streambooleandefault: false
If true, partial message deltas will be sent as server-sent events.
temperaturenumberdefault: 1
Sampling temperature between 0 and 2. Higher values make output more random.
top_pnumber
Nucleus sampling parameter. Must be greater than 0.
max_tokensinteger
Maximum number of tokens to generate.
max_completion_tokensinteger
Maximum number of completion tokens to generate.
presence_penaltynumberdefault: 0
Penalizes tokens based on presence in the text.
frequency_penaltynumberdefault: 0
Penalizes tokens based on frequency in the text.
seedinteger
Random seed for generation.
stoponeOf
A sequence, or a list of sequences, where the model stops generating.
logit_biasobject
Map of token id (as a string) to a bias from -100 to 100, as in OpenAI's API. Values outside that range are clamped.
toolsobject[]
Tools the model may call, in OpenAI's format ({"type": "function", "function": {...}}).
tool_choiceoneOf
auto, none, required, or a specific tool, as in OpenAI's API.
response_formatobject
Structured output: {"type": "json_object"} or {"type": "json_schema", "json_schema": {...}}, as in OpenAI's API.
ninteger
Not supported. Only one choice is returned; the parameter is ignored.
logprobsboolean
Not supported. The parameter is ignored.
top_logprobsinteger
Number of most likely tokens to return at each position. Requires logprobs: true.
top_kinteger
Sample from the k most likely tokens, for models that support it.
min_pnumber
Minimum token probability relative to the most likely token, for models that support it.
repetition_penaltynumber
Penalty for repeated tokens, for models that support it.
prompt_cache_keystring
Passed to providers that support prompt-cache routing keys. See caching.
sessionstring
An optional session id for grouping the requests of one conversation. Printable ASCII, up to 128 characters; default is reserved. Can also be sent in extra_body.
metadataobject
Request metadata, available to a saved router's conditional routing expressions and prompt templates. Can also be sent in extra_body.
reasoning_effortenum<string>
How much reasoning the model does, for models that support it. Case-insensitive. An unrecognized value is ignored, and for a model that does not support reasoning effort the parameter is dropped with a warning in metadata.attempts[].warnings. Takes precedence over extra_body.reasoning.effort.
Available options:noneminimallowmediumhighxhighmax
userstring
A stable id for your end user. When model is a saved router, it keeps a user on the same traffic-splitting variant; without it, the variant is chosen at random per request. It is not forwarded to the provider.
web_searchobject
Tool-based web search configuration. The LLM calls a search engine in a tool-calling loop, then synthesizes a grounded answer with url_citation annotations. Works with any LLM that supports tool calling. Mutually exclusive with web_search_options. See Web Search for details.
Show child attributes
engineenum<string>default: "exa"
Search backend: exa, brave or google, or native for the provider's built-in search.
Available options:exabravegooglenative
max_resultsintegerdefault: 3
Search results per search call, 1 to 10.
max_stepsintegerdefault: 1
Maximum search rounds before the final answer, 1 to 5.
web_search_optionsobject
Native web search using the provider's built-in search grounding (no tool loop). Supported by OpenAI, Azure OpenAI (search models only), Anthropic, Google / Vertex AI, and Groq. Mutually exclusive with web_search. See Web Search for details.
Show child attributes
search_context_sizeenum<string>
How much web context to retrieve. Defaults to medium. Used by OpenAI and Azure OpenAI; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.
Available options:lowmediumhigh
modalitiesenum<string>[]default: ["text"]
Output modalities to generate. Defaults to ["text"]. Include "image" to request image generation (e.g., ["text", "image"]). Currently supported for OpenAI and Google image models.
image_configobject
Configuration for image output. Optional when requesting image output via modalities: ["image"].
Show child attributes
aspect_ratiostring
Aspect ratio for the generated image (e.g., 1:1, 16:9, 9:16). Supported by Google models only.
image_sizestring
Size of the generated image.
- Google: model-specific sizes such as
1Kor2K. - OpenAI: pixel dimensions as WxH (e.g.,
1024x1024).
partial_imagesinteger
Number of partial/progressive image previews during streaming. Only used with stream: true. Defaults to 1 if unset. Supported by OpenAI models only.
ninteger
Number of images to generate.
- Google: Only support 1.
- OpenAI: Support 1-10.
extra_bodyobject
Optional parameters for model routing and optimization.
Show child attributes
modelsstring[]
Model ids for fallbacks, or the candidate pool for auto. A bare provider name (for example openai) stands for all of that provider's models. auto is not allowed here: at the top level it returns 400, and inside extra_body it matches no model.
ignorestring[]
Providers or models to exclude from selection. A provider/model named in model is never excluded; for a bare model id, ignore can exclude its providers. auto is not allowed here: at the top level it returns 400, and inside extra_body it matches no model.
sortstring[]
The sorting strategy for auto selection and fallbacks: price, latency, throughput, intelligence, math, coding. Add _asc or _desc to set the direction (for example latency_desc).
reasoningobject
Reasoning configuration. effort applies only when the top-level reasoning_effort is not set.
Show child attributes
effortenum<string>
Available options:noneminimallowmediumhighxhighmax
max_tokensinteger
Reasoning token budget, for models that take one.
excludeboolean
Leave the reasoning text out of the response.
providerobject
Provider routing configuration. Use when model is specified without a provider prefix (e.g., deepseek-v4-flash) to control which providers are tried and in what order.
Show child attributes
orderstring[]
Explicit list of providers to try, in order. Example: ["groq", "deepinfra"]. When specified, providers are tried in this exact order (sort criteria will be ignored), and providers not in the list are not used.
allow_fallbacksbooleandefault: true
Whether to allow falling back to the next provider if the current one fails. Defaults to true.
prompt_variablesobject
Variables for the prompt templates of a saved router (not for messages sent with each request). Values can be any JSON type. They are also available to conditional routing expressions, where metadata wins on a name clash. See prompt variables.
web_searchobject
For OpenAI SDK compatibility, pass web_search via extra_body (equivalent to setting it at the top level). Tool-based web search configuration. Mutually exclusive with web_search_options. See Web Search for details.
Show child attributes
engineenum<string>default: "exa"
Search backend: exa, brave or google, or native for the provider's built-in search.
Available options:exabravegooglenative
max_resultsintegerdefault: 3
Search results per search call, 1 to 10.
max_stepsintegerdefault: 1
Maximum search rounds before the final answer, 1 to 5.
web_search_optionsobject
For OpenAI SDK compatibility, pass web_search_options via extra_body (equivalent to setting it at the top level). Native web search grounding. Mutually exclusive with web_search. See Web Search for details.
Show child attributes
search_context_sizeenum<string>
How much web context to retrieve. Defaults to medium. Used by OpenAI and Azure OpenAI; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.
Available options:lowmediumhigh
fallbackobject
Fall back to the next model when the first token is slow.
Show child attributes
ttft_timeoutstringrequired
A duration of at least 300ms, for example "800ms" or "1.5s".
sessionstring
Same as the top-level session.
metadataobject
Same as the top-level metadata.
compressionobject
Prompt compression for this request. See prompt compression.
audioobject
Convert the generated text to Inworld TTS audio. Supports streaming and non-streaming requests. See LLM + TTS.
Show child attributes
voicestringrequired
Voice ID for speech synthesis.
modelstringrequired
Inworld TTS model ID; must start with inworld-. Independent of the top-level LLM model.
formatenum<string>default: "pcm16"
Encoding of the returned audio: pcm16 (16-bit signed little-endian PCM), mulaw (G.711 µ-law) or alaw (G.711 A-law). All are mono and headerless. Any other value returns 400.
Available options:pcm16mulawalaw
sample_rateinteger
Output sample rate in Hz. pcm16 accepts 8000, 16000, 22050, 24000, 32000, 44100 or 48000 (default 48000); mulaw and alaw accept 8000 only, the default. Any other value returns 400.
timestamp_typeenum<string>
Opt into timing entries for words or characters. Values are lowercase. Omit to disable timestamps. Alignment can add latency.
Available options:wordcharacter
idstring
Unique identifier for the chat completion.
objectstring
Object type, always 'chat.completion'.
createdinteger
Unix timestamp when the completion was created.
modelstring
The model that was actually used.
choicesobject[]
List of chat completion choices.
Show child attributes
indexinteger
messageobject
Show child attributes
rolestring
Always 'assistant' for responses.
contentstring
The generated text. An empty string when the model returned only tool calls, or when Inworld TTS audio is returned (read audio.transcript for the spoken text).
tool_callsobject[]
Tool calls generated by the model (when using tools).
Show child attributes
idstring
typestring
functionobject
Show child attributes
namestring
argumentsstring
audioobject
Synthesized Inworld TTS audio in a non-streaming completion. The message content is empty; the spoken text is in transcript.
Show child attributes
idstring
Identifier for the generated audio.
datastring
Base64-encoded mono, headerless audio in the requested audio.format, at the requested audio.sample_rate. Defaults: pcm16 (16-bit little-endian) at 48,000 Hz; mulaw and alaw are always 8,000 Hz.
transcriptstring
The synthesized text.
timestampsobject[]
Timing entries when audio.timestamp_type was requested and alignment is available. Times refer to the complete response audio.
Show child attributes
tokenstringrequired
Word or character associated with this timing entry.
start_timenumberrequired
Start time in seconds from the beginning of the complete response audio, including preceding sentences.
end_timenumberrequired
End time in seconds from the beginning of the complete response audio, including preceding sentences.
reasoningstring
The model's reasoning text, for reasoning models that return it.
annotationsobject[]
Citations, for example url_citation entries from web search.
finish_reasonenum<string>
Why the model stopped.
Available options:stoplengthtool_callscontent_filter
usageobject
Token usage statistics.
Show child attributes
prompt_tokensinteger
Tokens in the prompt.
completion_tokensinteger
Tokens in the completion.
total_tokensinteger
Total tokens used.
prompt_tokens_detailsobject
Always includes cached_tokens (prompt tokens read from cache). cache_write_tokens, audio_tokens, image_tokens and text_tokens appear when nonzero.
Show child attributes
cached_tokensinteger
cache_write_tokensinteger
completion_tokens_detailsobject
reasoning_tokens appears when nonzero.
Show child attributes
reasoning_tokensinteger
metadataobject
Routing metadata providing transparency into model selection decisions.
Show child attributes
attemptsobject[]
List of model attempts, including both successful and failed attempts.
Show child attributes
modelstring
The model identifier that was attempted.
successboolean
Whether this attempt succeeded.
time_to_first_token_msinteger
Time to receive the first token in milliseconds.
status_codeinteger
HTTP status of a failed attempt.
errorstring
Why the attempt failed.
duration_msnumber
Duration of the attempt in milliseconds.
warningsstring[]
Notes about the attempt, for example parameters skipped because the model does not support them.
generation_idstring
Unique identifier for tracing this request in the Inworld Portal.
reasoningstring
Human-readable explanation of why a model was selected based on the routing strategy.
total_duration_msinteger
Total request duration in milliseconds. Not included on streamed responses, where metadata arrives on the first chunk.
route_idstring
The saved router's route that was chosen, when model is a saved router.
variant_idstring
The saved router's variant that was chosen, when model is a saved router.
compression_warningsstring[]
Warnings from prompt compression, when it ran. Non-streaming responses only.
compressionobject
Stats from prompt compression, present when at least one message was compressed. Non-streaming responses only.
Show child attributes
original_tokensinteger
compressed_tokensinteger
saved_tokensinteger