Capabilities
Web search
Ground chat completions with web search
You can ground LLM responses with real-time web search results, either on a router variant or on a single request. Two modes are available:
| Mode | How it works | Supported models |
|---|---|---|
| Router-run | The LLM calls a search engine in a tool-calling loop, then synthesizes a final answer | Any LLM that supports tool calling |
| Native | The provider's built-in search grounding (no tool loop) | OpenAI, Azure OpenAI (your own key, search models only), Anthropic, Google AI Studio, Vertex AI, Groq |
On a router variant, both modes use the web_search field. On a chat completion request, you can also use web_search_options for native search.
Web search also works on the Responses API.
Router-run web search
Add a web_search object to your router variant. When a request is routed to that variant, the router injects a search tool, lets the LLM call it in a loop, and returns a grounded answer with url_citation annotations.
| Parameter | Type | Default | Description |
|---|---|---|---|
engine | string | exa | Options: exa, brave, google, native (uses the provider's built-in search; see Native web search) |
max_results | int | 3 | Number of search results returned per step (1–10) |
max_steps | int | 1 | Maximum tool-call rounds before final synthesis (1–5) |
An unknown engine or an out-of-range max_results or max_steps returns 400.
{
"variant_id": "search-grounded",
"model_id": "openai/gpt-4o",
"web_search": {
"engine": "exa",
"max_results": 5,
"max_steps": 2
}
}How it works:
- The router injects a search tool and sends the request to the LLM.
- The LLM calls the search tool with a query.
- The search engine returns results, which are injected back into the conversation.
- Steps 2–3 repeat up to
max_stepstimes. - The LLM synthesizes a final answer with
url_citationannotations.
Native web search
Native search uses the provider's built-in search grounding. It skips the tool-calling loop entirely, and the provider handles search internally.
On a router variant, set web_search with engine set to native:
{
"variant_id": "native-search",
"model_id": "openai/gpt-5.4",
"web_search": {
"engine": "native"
}
}On a chat completion request, you can instead pass web_search_options:
| Parameter | Type | Default | Description |
|---|---|---|---|
search_context_size | string | "medium" | Amount of search context: "low", "medium", or "high" |
search_context_size is the only supported field. Other fields, such as user_location, are ignored. It defaults to medium. OpenAI and Azure OpenAI use it directly; on Anthropic it sets how many searches the model can run; Google and Groq ignore it.
{
"model": "openai/gpt-5.4",
"messages": [{ "role": "user", "content": "What happened in tech news today?" }],
"web_search_options": {
"search_context_size": "high"
}
}Supported providers: OpenAI, Azure OpenAI (with your own Azure key, search models only), Anthropic, Google AI Studio / Vertex AI, and Groq. On OpenAI, native search works with any model, not only search models.
Per-request web search
You can also pass web_search or web_search_options directly on a chat completion request instead of configuring it on a variant. The fields and values are the same as above.
- Sending both
web_searchandweb_search_optionson the same request returns400. - Request-level settings take precedence over the variant's
web_search.
Both fields can be passed at the top level of the request body or inside extra_body for OpenAI SDK compatibility.
Citations & streaming
Assistant messages may include OpenAI-style annotations (e.g. type: "url_citation" with url, title, content).
With stream: true, annotations are delivered on the last SSE chunk, alongside finish_reason.