Get Started
Turn Detection
Detect end-of-turn automatically or control turn boundaries manually with the Inworld STT streaming API.
Turn detection identifies when a speaker has finished talking — the core signal a voice agent needs to know when to respond. The STT streaming API supports turn detection out of the box: the server detects end-of-turn automatically, and you can tune its sensitivity or take full manual control.
Turn detection is available on the WebSocket streaming endpoint. Sync (file upload) transcription processes complete audio files, so turn detection does not apply.
How it works
With inworld/inworld-stt-1 streaming, turn detection runs by default — no configuration required:
- As you stream audio, the server returns interim (partial) transcription results.
- When the server detects end-of-turn (for example, a sustained pause), it finalizes the transcript for that turn (
isFinal: true). - Speech after the turn boundary starts a new transcript.
A pause can split the transcript into two final results. The server waits for the configured minimum silence before checking for turn completion. It then ends the turn when the model is sufficiently confident, or when the maximum silence threshold is reached.
The server also emits voice-activity events you can use to drive application behavior (e.g., interrupt playback when the user starts speaking):
| Event | Meaning |
|---|---|
speechStarted | Voice activity detected in the audio stream |
speechStopped | Silence detected after speech has stopped |
Tuning automatic turn detection
Adjust sensitivity via transcribeConfig in the first WebSocket message:
| Field | Type | Default | Description |
|---|---|---|---|
endOfTurnConfidenceThreshold | float | 0.4 | Confidence required to declare end-of-turn. Higher values reduce false positives (fewer premature turn splits) at the cost of slower turn detection. Range: 0.0–1.0 |
inworldSttV1Config.minEndOfTurnSilenceWhenConfident | integer (ms) | — | Minimum trailing silence before checking for turn completion. Range: 20–5000 |
inworldSttV1Config.maxTurnSilence | integer (ms) | 300 | Trailing silence that ends the turn even if confidence is insufficient, once the minimum silence has been reached. Range: 20–5000 |
inworldSttV1Config.vadThreshold | float | 0.15 | Voice activity detection threshold. Range: 0.0–1.0 |
Server defaults can change. Set explicit values when tuning behavior; the values below are example settings. For application idle-stop behavior, maintain a client-side timer: inactivityTimeoutSeconds is not applied by inworld/inworld-stt-1.
{
"transcribeConfig": {
"modelId": "inworld/inworld-stt-1",
"audioEncoding": "LINEAR16",
"endOfTurnConfidenceThreshold": 0.7,
"inworldSttV1Config": {
"minEndOfTurnSilenceWhenConfident": 300,
"maxTurnSilence": 1200
}
}
}Keep minEndOfTurnSilenceWhenConfident less than or equal to maxTurnSilence. The minimum is checked first, so setting it higher than the maximum delays the maximum-silence fallback too. Set both fields when tuning how long the server waits through pauses.
Silence thresholds apply to the streamed audio and are evaluated as audio is processed. They do not guarantee transcript delivery or response latency.
Manual turn control
To hand turn control fully to the client, disable server-side voice activity detection by setting vadThreshold to 0:
{
"transcribeConfig": {
"modelId": "inworld/inworld-stt-1",
"audioEncoding": "LINEAR16",
"inworldSttV1Config": {
"vadThreshold": 0
}
}
}With VAD disabled, the server no longer splits turns automatically. Signal turn boundaries yourself:
- Send an
endTurnmessage at the end of each speaker turn to finalize the transcript. - When done sending audio, flush the final audio chunk and send
closeStreamonce. It finalizes pending audio, so do not sendendTurnimmediately before it. Keep receiving final results and usage until the server closes the stream.
With manual turn control, a single turn has a maximum length (currently around 30 seconds; subject to change). Send endTurn regularly at natural turn boundaries rather than relying on the limit.
Choosing a mode
| Mode | When to use |
|---|---|
| Automatic (default) | Voice agents and live transcription where the server should decide when the speaker is done |
| Automatic, tuned | Environments with background noise, slow speakers, or domain-specific pacing — adjust thresholds to reduce premature or delayed turn splits |
Manual (vadThreshold: 0) | Push-to-talk UIs, client-side VAD, or applications with their own turn-taking logic |