Capabilities
Vision and audio input
Send images and audio to models through LLM Router chat completions, and see how the router validates and routes them.
Vision lets a model read images that you send alongside text. Some models also accept audio input. On LLM Router you send both the same way: use the array form of a message's content, with one typed part per piece of content.
| Part type | Shape | Used for |
|---|---|---|
text | {"type": "text", "text": "..."} | Text. |
image_url | {"type": "image_url", "image_url": {"url": "...", "detail": "auto"}} | Images, as a URL or a base64 data URI. |
input_audio | {"type": "input_audio", "input_audio": {"data": "...", "format": "wav"}} | Base64-encoded audio. |
Whether a model accepts images or audio depends on the model. Refer to the provider's documentation for supported input types, formats, and limits.
Vision
Pass the image as an HTTP(S) URL or as a base64 data URI in image_url.url:
curl --request POST \
--url https://api.inworld.ai/v1/chat/completions \
--header "Authorization: Basic $INWORLD_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "google-ai-studio/gemini-2.5-flash",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Describe this image." },
{ "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQ..." } }
]
}
]
}'image_url.detail is optional and defaults to auto.
Image URLs on Google models
Google AI Studio and Vertex AI do not fetch images from HTTP(S) URLs. Send those models a base64 data URI.
The router checks this against the models a request can reach:
- If every candidate model is served by Google, an HTTP(S) image URL returns
400. - If the request can also reach a model from another provider, for example through fallbacks or
model: "auto", the router uses that model instead.
A data URI works on every provider, so prefer it when a request can route to more than one.
Audio input
Pass base64-encoded audio in input_audio.data and name its format:
curl --request POST \
--url https://api.inworld.ai/v1/chat/completions \
--header "Authorization: Basic $INWORLD_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "google-ai-studio/gemini-2.5-flash",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Transcribe this audio." },
{ "type": "input_audio", "input_audio": { "data": "UklGRi...", "format": "wav" } }
]
}
]
}'The router passes format to the provider without checking it, so the formats you can use are the ones the model accepts. Anthropic models do not take audio input: the audio part is left out of the request and the model sees only the rest of the message.
To transcribe speech, Inworld STT is the dedicated API. To hold a spoken conversation with a model, use the Realtime API.
Validation
- A data URI in
image_urlmust have animage/*media type, and one ininput_audioanaudio/*media type. Anything else returns400. - Video input is not supported and returns
400. - Content parts of any other type, such as
file, are skipped rather than rejected. The request can succeed without the model having seen that part, so check that every part you send is one of the three types above.
How routing uses input types
- With
model: "auto", the router only considers models that accept the input types in your request. - With a named model, that model is always tried, and only its fallbacks are filtered by input type.
Related
- Voice responses return spoken audio from a chat completion.
- Chat completions covers the request basics.
- API reference lists every parameter.