Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Chat Completions

Create chat completion

Generate a response for the given chat conversation

POST/v1/chat/completions

Call 100+ models from various providers directly through our unified API, or set model to auto for automatic model selection based on criteria like price, latency, or performance.

For more advanced routing — such as conditional routing, A/B testing across variants, and reusable configurations — create a router and reference it via the model field (e.g., inworld/my-router).

For web-grounded answers, use extra_body.web_search.

For voice output, see Voice responses, including optional word or character timestamps.

Authorizations

Authorizationstringrequired

Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.

Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Body

application/json

modelstringrequired

The model to use, which can be:

  • A model id (e.g., deepseek-v4-flash). The best provider is automatically selected by latency, or you can control provider selection via extra_body.provider. See Models for available models.
  • A provider-prefixed model id (e.g., openai/gpt-5.4). This specifies the provider and model to use.
  • auto for automatic model selection based on criteria like price, latency, or intelligence
  • A router, which is specified by inworld/<router-name>. The router name must be prefixed by inworld/.

messagesobject[]required

A list of messages comprising the conversation so far.

If using a router where a prompt is specified, these messages will be appended to the prompt.

Show child attributes

roleenum<string>required

The role of the message author.

Available options:systemuserassistanttool

contentoneOf

The content of the message. Can be a string for text-only messages, an array of content parts for multimodal messages, or null for assistant messages with tool_calls.

tool_callsobject[]

Tool calls generated by the model (assistant messages only).

Show child attributes

idstringrequired

ID of the tool call.

typeenum<string>required

The type of the tool call. Always 'function'.

Available options:function

functionobjectrequired

The function that the model called.

Show child attributes

namestringrequired

The name of the function to call.

argumentsstringrequired

The arguments to call the function with, as a JSON string.

tool_call_idstring

Tool call ID this message is responding to (tool role only).

streambooleandefault: false

If true, partial message deltas will be sent as server-sent events.

temperaturenumberdefault: 1

Sampling temperature between 0 and 2. Higher values make output more random.

top_pnumber

Nucleus sampling parameter. Must be greater than 0.

max_tokensinteger

Maximum number of tokens to generate.

max_completion_tokensinteger

Maximum number of completion tokens to generate.

presence_penaltynumberdefault: 0

Penalizes tokens based on presence in the text.

frequency_penaltynumberdefault: 0

Penalizes tokens based on frequency in the text.

seedinteger

Random seed for generation.

stoponeOf

A sequence, or a list of sequences, where the model stops generating.

logit_biasobject

Map of token id (as a string) to a bias from -100 to 100, as in OpenAI's API. Values outside that range are clamped.

toolsobject[]

Tools the model may call, in OpenAI's format ({"type": "function", "function": {...}}).

tool_choiceoneOf

auto, none, required, or a specific tool, as in OpenAI's API.

response_formatobject

Structured output: {"type": "json_object"} or {"type": "json_schema", "json_schema": {...}}, as in OpenAI's API.

ninteger

Not supported. Only one choice is returned; the parameter is ignored.

logprobsboolean

Not supported. The parameter is ignored.

top_logprobsinteger

Number of most likely tokens to return at each position. Requires logprobs: true.

top_kinteger

Sample from the k most likely tokens, for models that support it.

min_pnumber

Minimum token probability relative to the most likely token, for models that support it.

repetition_penaltynumber

Penalty for repeated tokens, for models that support it.

prompt_cache_keystring

Passed to providers that support prompt-cache routing keys. See caching.

sessionstring

An optional session id for grouping the requests of one conversation. Printable ASCII, up to 128 characters; default is reserved. Can also be sent in extra_body.

metadataobject

Request metadata, available to a saved router's conditional routing expressions and prompt templates. Can also be sent in extra_body.

reasoning_effortenum<string>

How much reasoning the model does, for models that support it. Case-insensitive. An unrecognized value is ignored, and for a model that does not support reasoning effort the parameter is dropped with a warning in metadata.attempts[].warnings. Takes precedence over extra_body.reasoning.effort.

Available options:noneminimallowmediumhighxhighmax

userstring

A stable id for your end user. When model is a saved router, it keeps a user on the same traffic-splitting variant; without it, the variant is chosen at random per request. It is not forwarded to the provider.

web_searchobject

Tool-based web search configuration. The LLM calls a search engine in a tool-calling loop, then synthesizes a grounded answer with url_citation annotations. Works with any LLM that supports tool calling. Mutually exclusive with web_search_options. See Web Search for details.

Show child attributes

engineenum<string>default: "exa"

Search backend: exa, brave or google, or native for the provider's built-in search.

Available options:exabravegooglenative

max_resultsintegerdefault: 3

Search results per search call, 1 to 10.

max_stepsintegerdefault: 1

Maximum search rounds before the final answer, 1 to 5.

web_search_optionsobject

Native web search using the provider's built-in search grounding (no tool loop). Supported by OpenAI, Azure OpenAI (search models only), Anthropic, Google / Vertex AI, and Groq. Mutually exclusive with web_search. See Web Search for details.

Show child attributes

search_context_sizeenum<string>

How much web context to retrieve. Defaults to medium. Used by OpenAI and Azure OpenAI; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.

Available options:lowmediumhigh

modalitiesenum<string>[]default: ["text"]

Output modalities to generate. Defaults to ["text"]. Include "image" to request image generation (e.g., ["text", "image"]). Currently supported for OpenAI and Google image models.

image_configobject

Configuration for image output. Optional when requesting image output via modalities: ["image"].

Show child attributes

aspect_ratiostring

Aspect ratio for the generated image (e.g., 1:1, 16:9, 9:16). Supported by Google models only.

image_sizestring

Size of the generated image.

  • Google: model-specific sizes such as 1K or 2K.
  • OpenAI: pixel dimensions as WxH (e.g., 1024x1024).

partial_imagesinteger

Number of partial/progressive image previews during streaming. Only used with stream: true. Defaults to 1 if unset. Supported by OpenAI models only.

ninteger

Number of images to generate.

  • Google: Only support 1.
  • OpenAI: Support 1-10.

extra_bodyobject

Optional parameters for model routing and optimization.

Show child attributes

modelsstring[]

Model ids for fallbacks, or the candidate pool for auto. A bare provider name (for example openai) stands for all of that provider's models. auto is not allowed here: at the top level it returns 400, and inside extra_body it matches no model.

ignorestring[]

Providers or models to exclude from selection. A provider/model named in model is never excluded; for a bare model id, ignore can exclude its providers. auto is not allowed here: at the top level it returns 400, and inside extra_body it matches no model.

sortstring[]

The sorting strategy for auto selection and fallbacks: price, latency, throughput, intelligence, math, coding. Add _asc or _desc to set the direction (for example latency_desc).

reasoningobject

Reasoning configuration. effort applies only when the top-level reasoning_effort is not set.

Show child attributes

effortenum<string>

Available options:noneminimallowmediumhighxhighmax

max_tokensinteger

Reasoning token budget, for models that take one.

excludeboolean

Leave the reasoning text out of the response.

providerobject

Provider routing configuration. Use when model is specified without a provider prefix (e.g., deepseek-v4-flash) to control which providers are tried and in what order.

Show child attributes

orderstring[]

Explicit list of providers to try, in order. Example: ["groq", "deepinfra"]. When specified, providers are tried in this exact order (sort criteria will be ignored), and providers not in the list are not used.

allow_fallbacksbooleandefault: true

Whether to allow falling back to the next provider if the current one fails. Defaults to true.

prompt_variablesobject

Variables for the prompt templates of a saved router (not for messages sent with each request). Values can be any JSON type. They are also available to conditional routing expressions, where metadata wins on a name clash. See prompt variables.

web_searchobject

For OpenAI SDK compatibility, pass web_search via extra_body (equivalent to setting it at the top level). Tool-based web search configuration. Mutually exclusive with web_search_options. See Web Search for details.

Show child attributes

engineenum<string>default: "exa"

Search backend: exa, brave or google, or native for the provider's built-in search.

Available options:exabravegooglenative

max_resultsintegerdefault: 3

Search results per search call, 1 to 10.

max_stepsintegerdefault: 1

Maximum search rounds before the final answer, 1 to 5.

web_search_optionsobject

For OpenAI SDK compatibility, pass web_search_options via extra_body (equivalent to setting it at the top level). Native web search grounding. Mutually exclusive with web_search. See Web Search for details.

Show child attributes

search_context_sizeenum<string>

How much web context to retrieve. Defaults to medium. Used by OpenAI and Azure OpenAI; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.

Available options:lowmediumhigh

fallbackobject

Fall back to the next model when the first token is slow.

Show child attributes

ttft_timeoutstringrequired

A duration of at least 300ms, for example "800ms" or "1.5s".

sessionstring

Same as the top-level session.

metadataobject

Same as the top-level metadata.

compressionobject

Prompt compression for this request. See prompt compression.

audioobject

Convert the generated text to Inworld TTS audio. Supports streaming and non-streaming requests. See LLM + TTS.

Show child attributes

voicestringrequired

Voice ID for speech synthesis.

modelstringrequired

Inworld TTS model ID; must start with inworld-. Independent of the top-level LLM model.

formatenum<string>default: "pcm16"

Encoding of the returned audio: pcm16 (16-bit signed little-endian PCM), mulaw (G.711 µ-law) or alaw (G.711 A-law). All are mono and headerless. Any other value returns 400.

Available options:pcm16mulawalaw

sample_rateinteger

Output sample rate in Hz. pcm16 accepts 8000, 16000, 22050, 24000, 32000, 44100 or 48000 (default 48000); mulaw and alaw accept 8000 only, the default. Any other value returns 400.

timestamp_typeenum<string>

Opt into timing entries for words or characters. Values are lowercase. Omit to disable timestamps. Alignment can add latency.

Available options:wordcharacter

Response

200 - application/json

idstring

Unique identifier for the chat completion.

objectstring

Object type, always 'chat.completion'.

createdinteger

Unix timestamp when the completion was created.

modelstring

The model that was actually used.

choicesobject[]

List of chat completion choices.

Show child attributes

indexinteger

messageobject

Show child attributes

rolestring

Always 'assistant' for responses.

contentstring

The generated text. An empty string when the model returned only tool calls, or when Inworld TTS audio is returned (read audio.transcript for the spoken text).

tool_callsobject[]

Tool calls generated by the model (when using tools).

Show child attributes

idstring

typestring

functionobject

Show child attributes

namestring

argumentsstring

audioobject

Synthesized Inworld TTS audio in a non-streaming completion. The message content is empty; the spoken text is in transcript.

Show child attributes

idstring

Identifier for the generated audio.

datastring

Base64-encoded mono, headerless audio in the requested audio.format, at the requested audio.sample_rate. Defaults: pcm16 (16-bit little-endian) at 48,000 Hz; mulaw and alaw are always 8,000 Hz.

transcriptstring

The synthesized text.

timestampsobject[]

Timing entries when audio.timestamp_type was requested and alignment is available. Times refer to the complete response audio.

Show child attributes

tokenstringrequired

Word or character associated with this timing entry.

start_timenumberrequired

Start time in seconds from the beginning of the complete response audio, including preceding sentences.

end_timenumberrequired

End time in seconds from the beginning of the complete response audio, including preceding sentences.

reasoningstring

The model's reasoning text, for reasoning models that return it.

annotationsobject[]

Citations, for example url_citation entries from web search.

finish_reasonenum<string>

Why the model stopped.

Available options:stoplengthtool_callscontent_filter

usageobject

Token usage statistics.

Show child attributes

prompt_tokensinteger

Tokens in the prompt.

completion_tokensinteger

Tokens in the completion.

total_tokensinteger

Total tokens used.

prompt_tokens_detailsobject

Always includes cached_tokens (prompt tokens read from cache). cache_write_tokens, audio_tokens, image_tokens and text_tokens appear when nonzero.

Show child attributes

cached_tokensinteger

cache_write_tokensinteger

completion_tokens_detailsobject

reasoning_tokens appears when nonzero.

Show child attributes

reasoning_tokensinteger

metadataobject

Routing metadata providing transparency into model selection decisions.

Show child attributes

attemptsobject[]

List of model attempts, including both successful and failed attempts.

Show child attributes

modelstring

The model identifier that was attempted.

successboolean

Whether this attempt succeeded.

time_to_first_token_msinteger

Time to receive the first token in milliseconds.

status_codeinteger

HTTP status of a failed attempt.

errorstring

Why the attempt failed.

duration_msnumber

Duration of the attempt in milliseconds.

warningsstring[]

Notes about the attempt, for example parameters skipped because the model does not support them.

generation_idstring

Unique identifier for tracing this request in the Inworld Portal.

reasoningstring

Human-readable explanation of why a model was selected based on the routing strategy.

total_duration_msinteger

Total request duration in milliseconds. Not included on streamed responses, where metadata arrives on the first chunk.

route_idstring

The saved router's route that was chosen, when model is a saved router.

variant_idstring

The saved router's variant that was chosen, when model is a saved router.

compression_warningsstring[]

Warnings from prompt compression, when it ran. Non-streaming responses only.

compressionobject

Stats from prompt compression, present when at least one message was compressed. Non-streaming responses only.

Show child attributes

original_tokensinteger

compressed_tokensinteger

saved_tokensinteger