Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Web search

Ground chat completions with web search

You can ground LLM responses with real-time web search results, either on a router variant or on a single request. Two modes are available:

ModeHow it worksSupported models
Router-runThe LLM calls a search engine in a tool-calling loop, then synthesizes a final answerAny LLM that supports tool calling
NativeThe provider's built-in search grounding (no tool loop)OpenAI, Azure OpenAI (your own key, search models only), Anthropic, Google AI Studio, Vertex AI, Groq

On a router variant, both modes use the web_search field. On a chat completion request, you can also use web_search_options for native search.

Web search also works on the Responses API.

Add a web_search object to your router variant. When a request is routed to that variant, the router injects a search tool, lets the LLM call it in a loop, and returns a grounded answer with url_citation annotations.

ParameterTypeDefaultDescription
enginestringexaOptions: exa, brave, google, native (uses the provider's built-in search; see Native web search)
max_resultsint3Number of search results returned per step (1–10)
max_stepsint1Maximum tool-call rounds before final synthesis (1–5)

An unknown engine or an out-of-range max_results or max_steps returns 400.

Variant configuration
{
  "variant_id": "search-grounded",
  "model_id": "openai/gpt-4o",
  "web_search": {
    "engine": "exa",
    "max_results": 5,
    "max_steps": 2
  }
}

How it works:

  1. The router injects a search tool and sends the request to the LLM.
  2. The LLM calls the search tool with a query.
  3. The search engine returns results, which are injected back into the conversation.
  4. Steps 2–3 repeat up to max_steps times.
  5. The LLM synthesizes a final answer with url_citation annotations.

Native search uses the provider's built-in search grounding. It skips the tool-calling loop entirely, and the provider handles search internally.

On a router variant, set web_search with engine set to native:

Variant configuration
{
  "variant_id": "native-search",
  "model_id": "openai/gpt-5.4",
  "web_search": {
    "engine": "native"
  }
}

On a chat completion request, you can instead pass web_search_options:

ParameterTypeDefaultDescription
search_context_sizestring"medium"Amount of search context: "low", "medium", or "high"

search_context_size is the only supported field. Other fields, such as user_location, are ignored. It defaults to medium. OpenAI and Azure OpenAI use it directly; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.

Request body
{
  "model": "openai/gpt-5.4",
  "messages": [{ "role": "user", "content": "What happened in tech news today?" }],
  "web_search_options": {
    "search_context_size": "high"
  }
}

Supported providers: OpenAI, Azure OpenAI (with your own Azure key, search models only), Anthropic, Google AI Studio / Vertex AI, and Groq. On OpenAI, native search works with any model, not only search models.

You can also pass web_search or web_search_options directly on a chat completion request instead of configuring it on a variant. The fields and values are the same as above.

  • Sending both web_search and web_search_options on the same request returns 400.
  • Request-level settings take precedence over the variant's web_search.

Both fields can be passed at the top level of the request body or inside extra_body for OpenAI SDK compatibility.

Citations & streaming

Assistant messages may include OpenAI-style annotations (e.g. type: "url_citation" with url, title, content).

With stream: true, annotations are delivered on the last SSE chunk, alongside finish_reason.