Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Vision and audio input

Send images and audio to models through LLM Router chat completions, and see how the router validates and routes them.

Vision lets a model read images that you send alongside text. Some models also accept audio input. On LLM Router you send both the same way: use the array form of a message's content, with one typed part per piece of content.

Part typeShapeUsed for
text{"type": "text", "text": "..."}Text.
image_url{"type": "image_url", "image_url": {"url": "...", "detail": "auto"}}Images, as a URL or a base64 data URI.
input_audio{"type": "input_audio", "input_audio": {"data": "...", "format": "wav"}}Base64-encoded audio.

Whether a model accepts images or audio depends on the model. Refer to the provider's documentation for supported input types, formats, and limits.

Vision

Pass the image as an HTTP(S) URL or as a base64 data URI in image_url.url:

bash
curl --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "google-ai-studio/gemini-2.5-flash",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "Describe this image." },
          { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQ..." } }
        ]
      }
    ]
  }'

image_url.detail is optional and defaults to auto.

Image URLs on Google models

Google AI Studio and Vertex AI do not fetch images from HTTP(S) URLs. Send those models a base64 data URI.

The router checks this against the models a request can reach:

  • If every candidate model is served by Google, an HTTP(S) image URL returns 400.
  • If the request can also reach a model from another provider, for example through fallbacks or model: "auto", the router uses that model instead.

A data URI works on every provider, so prefer it when a request can route to more than one.

Audio input

Pass base64-encoded audio in input_audio.data and name its format:

bash
curl --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "google-ai-studio/gemini-2.5-flash",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "Transcribe this audio." },
          { "type": "input_audio", "input_audio": { "data": "UklGRi...", "format": "wav" } }
        ]
      }
    ]
  }'

The router passes format to the provider without checking it, so the formats you can use are the ones the model accepts. Anthropic models do not take audio input: the audio part is left out of the request and the model sees only the rest of the message.

To transcribe speech, Inworld STT is the dedicated API. To hold a spoken conversation with a model, use the Realtime API.

Validation

  • A data URI in image_url must have an image/* media type, and one in input_audio an audio/* media type. Anything else returns 400.
  • Video input is not supported and returns 400.
  • Content parts of any other type, such as file, are skipped rather than rejected. The request can succeed without the model having seen that part, so check that every part you send is one of the three types above.

How routing uses input types

  • With model: "auto", the router only considers models that accept the input types in your request.
  • With a named model, that model is always tried, and only its fallbacks are filtered by input type.