Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Billing & usage

Performance

Monitor latency, errors, uptime, and rate limiting for your Inworld API traffic in Portal.

The Performance page shows how fast your requests are, how often they fail, and how often they hit your plan's limits. It covers the current workspace only.

Open it from Manage > Performance in Portal.

Performance is available on the Builder plan and above. To upgrade, see Billing.

Choose a product and time range

The tabs at the top switch between LLM Router, Text-to-Speech, Speech-to-Text and Realtime. Each tab shows metrics for that product only.

Pick Last 24 hours, Last 7 days, Last 30 days, or a custom range. The range is set when you open the page, so reload it to see the latest data.

Metric tabs

TabWhat it shows
OverviewKey metrics per model in one table
Latencyp50, p95 and p99 latency over time
ThroughputOutput tokens per second (LLM Router only)
CachingShare of prompt tokens read from the provider's prompt cache (LLM Router only)
ErrorsShare of requests that failed, split into client errors, server errors, rate limiting, and responses cut off mid-stream
UptimeShare of requests that did not fail with a server error
ThrottlingRequests rejected by rate limits or concurrency limits
ConcurrencyPeak concurrent connections

On the chart tabs, switch from Total to By model to compare up to eight models. Concurrency and Realtime show totals only.

p50 is the median: half of your requests were faster than this value. p95 and p99 show the slowest 5% and 1%. Percentiles are estimated from latency ranges, so treat them as approximate.

Overview table

The Overview tab lists one row per model, sorted by request count. For the LLM Router, the columns are:

ColumnDescription
ModelThe model that handled the requests
RequestsNumber of requests in the selected period
Throughput (p50)Median output tokens per second
TTFT (p50)Median time to first token
Latency (p50)Median time to complete a request
UptimeShare of requests that did not fail with a server error
Error rateShare of requests that failed with a client error, server error or rate limit, or whose streamed response was cut off before it finished
Cache utilShare of prompt tokens read from the prompt cache, shown when there is cache data
429sRequests that reached the model and got a 429 rate-limit response

Other products show their own latency columns:

ProductLatency columns
Text-to-SpeechFirst audio (p50), Synthesis (p50)
Speech-to-TextRecognition (p50) for Inworld STT models only, First transcript (p50)
RealtimePerceived text (p50), Perceived audio (p50). Each row is an STT, LLM and TTS model combination

Hover over an underlined column name to see its definition.

Reduce rate limiting

Requests rejected by your plan's rate or concurrency limits never reach a model, so they are not in the table. When there are any, a summary below the table shows how many were rejected and their share of all requests. The Errors and Throttling tabs include them in the Total view.

To raise these limits, upgrade your plan. See Rate limits for how limits work and Billing for the limits by plan.

Next steps