Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Get Started

Turn Detection

Detect end-of-turn automatically or control turn boundaries manually with the Inworld STT streaming API.

Turn detection identifies when a speaker has finished talking — the core signal a voice agent needs to know when to respond. The STT streaming API supports turn detection out of the box: the server detects end-of-turn automatically, and you can tune its sensitivity or take full manual control.

Turn detection is available on the WebSocket streaming endpoint. Sync (file upload) transcription processes complete audio files, so turn detection does not apply.

How it works

With inworld/inworld-stt-1 streaming, turn detection runs by default — no configuration required:

  1. As you stream audio, the server returns interim (partial) transcription results.
  2. When the server detects end-of-turn (for example, a sustained pause), it finalizes the transcript for that turn (isFinal: true).
  3. Speech after the turn boundary starts a new transcript.

A pause can split the transcript into two final results. The server waits for the configured minimum silence before checking for turn completion. It then ends the turn when the model is sufficiently confident, or when the maximum silence threshold is reached.

The server also emits voice-activity events you can use to drive application behavior (e.g., interrupt playback when the user starts speaking):

EventMeaning
speechStartedVoice activity detected in the audio stream
speechStoppedSilence detected after speech has stopped

Tuning automatic turn detection

Adjust sensitivity via transcribeConfig in the first WebSocket message:

FieldTypeDefaultDescription
endOfTurnConfidenceThresholdfloat0.4Confidence required to declare end-of-turn. Higher values reduce false positives (fewer premature turn splits) at the cost of slower turn detection. Range: 0.0–1.0
inworldSttV1Config.minEndOfTurnSilenceWhenConfidentinteger (ms)—Minimum trailing silence before checking for turn completion. Range: 20–5000
inworldSttV1Config.maxTurnSilenceinteger (ms)300Trailing silence that ends the turn even if confidence is insufficient, once the minimum silence has been reached. Range: 20–5000
inworldSttV1Config.vadThresholdfloat0.15Voice activity detection threshold. Range: 0.0–1.0

Server defaults can change. Set explicit values when tuning behavior; the values below are example settings. For application idle-stop behavior, maintain a client-side timer: inactivityTimeoutSeconds is not applied by inworld/inworld-stt-1.

json
{
  "transcribeConfig": {
    "modelId": "inworld/inworld-stt-1",
    "audioEncoding": "LINEAR16",
    "endOfTurnConfidenceThreshold": 0.7,
    "inworldSttV1Config": {
      "minEndOfTurnSilenceWhenConfident": 300,
      "maxTurnSilence": 1200
    }
  }
}

Keep minEndOfTurnSilenceWhenConfident less than or equal to maxTurnSilence. The minimum is checked first, so setting it higher than the maximum delays the maximum-silence fallback too. Set both fields when tuning how long the server waits through pauses.

Silence thresholds apply to the streamed audio and are evaluated as audio is processed. They do not guarantee transcript delivery or response latency.

Manual turn control

To hand turn control fully to the client, disable server-side voice activity detection by setting vadThreshold to 0:

json
{
  "transcribeConfig": {
    "modelId": "inworld/inworld-stt-1",
    "audioEncoding": "LINEAR16",
    "inworldSttV1Config": {
      "vadThreshold": 0
    }
  }
}

With VAD disabled, the server no longer splits turns automatically. Signal turn boundaries yourself:

  • Send an endTurn message at the end of each speaker turn to finalize the transcript.
  • When done sending audio, flush the final audio chunk and send closeStream once. It finalizes pending audio, so do not send endTurn immediately before it. Keep receiving final results and usage until the server closes the stream.

With manual turn control, a single turn has a maximum length (currently around 30 seconds; subject to change). Send endTurn regularly at natural turn boundaries rather than relying on the limit.

Choosing a mode

ModeWhen to use
Automatic (default)Voice agents and live transcription where the server should decide when the speaker is done
Automatic, tunedEnvironments with background noise, slow speakers, or domain-specific pacing — adjust thresholds to reduce premature or delayed turn splits
Manual (vadThreshold: 0)Push-to-talk UIs, client-side VAD, or applications with their own turn-taking logic

Next steps