Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Tool calling

Let models call your functions through LLM Router: define tools once in the OpenAI format, handle tool calls, and return results, on any provider.

Tool calling (also called function calling) lets a model call functions that you define. You describe your tools in the OpenAI format, and LLM Router translates them for the provider that serves the request, so the same request works on OpenAI, Anthropic, Google and other models, and keeps working when the router falls back to another provider.

Building a voice agent on a realtime session instead? See Tool calling for the Realtime API. Tools also work on the Responses API and the Anthropic Messages API.

How it works

  1. Send a chat completion with a tools array that describes your functions.
  2. If the model decides to call one, the response has finish_reason: "tool_calls" and a tool_calls array on the assistant message.
  3. Run the function in your own code.
  4. Send a second request with the assistant message and a role: "tool" message that carries the result. The model uses it to write the final answer.

The model never runs anything itself. Your application executes every tool call.

Tool calling needs a billing method on your workspace. Without one, a request that includes tools returns 400.

Define tools and get a tool call

cURL
curl --request POST \
  --url https://api.inworld.ai/v1/chat/completions \
  --header "Authorization: Basic $INWORLD_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "openai/gpt-5",
    "messages": [
      {"role": "user", "content": "What is the weather in Paris?"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_weather",
          "description": "Get the current weather for a city.",
          "parameters": {
            "type": "object",
            "properties": {
              "city": {"type": "string", "description": "City name, e.g. Paris"}
            },
            "required": ["city"]
          }
        }
      }
    ]
  }'

When the model calls the tool, the assistant message looks like this:

json
{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "id": "call_weather",
      "type": "function",
      "index": 0,
      "function": {
        "name": "get_weather",
        "arguments": "{\"city\": \"Paris\"}"
      }
    }
  ]
}

function.arguments is a JSON string, so parse it before use. content is an empty string rather than null when the model only calls tools.

Tool definition fields

FieldTypeDescription
typestring"function".
function.namestringThe name the model uses to call the tool.
function.descriptionstringWhat the tool does and when to use it. The model relies on this to decide whether to call the tool.
function.parametersobjectA JSON Schema object that describes the arguments. Any other JSON type returns 400.

Return the tool result

Append the assistant message exactly as you received it, then add one role: "tool" message per tool call with the matching tool_call_id:

Python
assistant_message = response.choices[0].message
messages.append(assistant_message)

for tool_call in assistant_message.tool_calls:
    args = json.loads(tool_call.function.arguments)
    result = get_weather(args["city"])  # your function
    messages.append({
        "role": "tool",
        "tool_call_id": tool_call.id,
        "content": json.dumps(result),
    })

final = client.chat.completions.create(
    model="openai/gpt-5",
    messages=messages,
    tools=tools,
)
print(final.choices[0].message.content)

The tool message content can be a string or an array of text parts:

json
{
  "role": "tool",
  "tool_call_id": "call_weather",
  "content": [{"type": "text", "text": "The temperature in Paris is 18ยฐC."}]
}

A tool message without a tool_call_id fails with 400.

A model can return several tool calls in one turn. Run each one and send back a tool message for every id before you ask for the next completion.

Send the assistant message back unchanged. Gemini models attach a thought_signature to each tool call's function object, and it has to come back on the follow-up request.

Control when tools are called

Set tool_choice to steer the model. The router translates each value to the provider's equivalent.

ValueBehavior
"auto"The model decides whether to call a tool.
"none"The model does not call tools.
"required"The model must call at least one tool.
{"type": "function", "function": {"name": "get_weather"}}The model must call the named tool.

Streaming

With stream: true, tool calls arrive as delta.tool_calls chunks. Each entry has an index that identifies the tool call, and the arguments arrive as string fragments:

json
{"index": 0, "id": "call_weather", "type": "function", "function": {"name": "get_weather", "arguments": ""}}
{"index": 0, "id": "", "type": "function", "function": {"name": "", "arguments": "{\"city\":"}}
{"index": 0, "id": "", "type": "function", "function": {"name": "", "arguments": " \"Paris\"}"}}

Group the chunks by index and concatenate function.arguments. Later chunks carry empty id and name strings instead of omitting them, so keep the first non-empty value of each. The final chunk has finish_reason: "tool_calls". See Streaming for the stream format.

Differences from OpenAI

The request and response shapes follow OpenAI's chat completions API, with these differences:

  • parallel_tool_calls is ignored. Models that can return several tool calls in one turn still do.
  • function.strict is ignored, so arguments are not guaranteed to match your schema. Validate them before you run the tool.
  • The legacy functions and function_call parameters are ignored. Use tools and tool_choice.
  • Object forms of tool_choice other than the named-function form are ignored.
  • Each entry in a non-streaming tool_calls array also has an index.

Tools with other features

  • Voice responses: with voice responses, a turn that calls a tool produces no audio, and the answer after the tool result is spoken.
  • Web search: router-run web search runs its own tool loop on the server, so you do not define or execute a search tool yourself.
  • Routing: tool definitions are translated per provider, so fallbacks and conditional routing work without changes to your tools.

Next steps