WebSocket
Transcribe audio (WebSocket)
Connect and authenticate
Connect to wss://api.inworld.ai/stt/v1/transcribe:streamBidirectional.
Server-side clients send Authorization: Basic <INWORLD_API_KEY> on the WebSocket
upgrade request. Use the Base64 credentials copied from the Portal verbatim; do
not Base64-encode them again.
const WebSocket = require('ws');
const ws = new WebSocket(
'wss://api.inworld.ai/stt/v1/transcribe:streamBidirectional',
{ headers: { Authorization: `Basic ${process.env.INWORLD_API_KEY}` } }
);This uses the Node.js ws package. For browser or mobile clients, keep the API
key on your server. Either relay audio through an authenticated backend
connection or connect directly using a one-time token
minted by your backend.
The browser's native WebSocket constructor cannot set an Authorization header.
For a direct connection, pass the token in the bearer_ WebSocket subprotocol:
// Your authenticated backend endpoint mints the token; never expose the API key.
// See the one-time-token guide for the backend handler and user authorization.
const response = await fetch('/api/inworld-token', {
method: 'POST',
credentials: 'same-origin',
cache: 'no-store',
});
if (!response.ok) throw new Error(`Token request failed (${response.status})`);
const { accessToken } = await response.json();
const ws = new WebSocket(
'wss://api.inworld.ai/stt/v1/transcribe:streamBidirectional',
['bearer_' + accessToken],
);Mint a fresh token for every connection attempt, including retries and reconnects.
Do not put tokens in URL query parameters. Clients that can set headers may send
Authorization: Bearer <accessToken> instead. Follow your application's CSRF
requirements when calling the backend token endpoint. See the
one-time-token guide
for the complete backend flow.
Handle connection/upgrade errors as well as JSON error messages after connection. A failed handshake may never produce a transcription or JSON error frame; report it to the caller instead of waiting indefinitely for a transcript.
Streaming message sequence
Send each client message below as a separate JSON text frame. Responses can arrive while you are sending audio; read them concurrently.
-
Configure the stream first:
json {"transcribeConfig":{"modelId":"inworld/inworld-stt-1","audioEncoding":"LINEAR16","sampleRateHertz":16000,"numberOfChannels":1}} -
Send audio chunks. Replace the placeholder with Base64-encoded raw PCM bytes:
json {"audioChunk":{"content":"<base64-encoded PCM>"}} -
Receive interim transcripts automatically. Each interim replaces the previous interim for the current turn; it is not a text delta:
json {"result":{"transcription":{"transcript":"Please open","isFinal":false}}} {"result":{"transcription":{"transcript":"Please open the door","isFinal":false}}} -
Commit each final once, then clear the interim. Later audio starts a new turn:
json {"result":{"transcription":{"transcript":"Please open the door.","isFinal":true}}} -
To finish sending audio, flush your last partial audio chunk and send:
json {"closeStream":{}}
Keep reading until the server closes the stream, including any final transcripts
and result.usage. A final transcript ends a turn, not necessarily the stream.
Do not close the socket immediately after sending the control message.
Turn boundaries and shutdown
| Client message | Use |
|---|---|
{"endTurn":{}} | Finalize the current turn while keeping the session open for more audio |
{"closeStream":{}} | Finish audio input, finalize pending audio, and let the server drain and close the session |
At shutdown, send closeStream once; do not precede it with endTurn for the same
audio. Repeated finalization of an empty or silent turn can produce unwanted
transcripts. With automatic turn detection, pauses can already produce finals;
listen for them without sending an extra endTurn for each pause.
endTurn alone does not finish the session or request final usage. Disconnecting
the transport can lose pending results and usage. Set a client-side shutdown
timeout and surface a timeout as incomplete transcription, rather than silently
treating a trailing interim as final. Result cadence and finalization latency
vary with audio and model processing; a short quiet period is not a completion
signal.
Voice activity events
Speech events are separate from transcription results and do not themselves finalize a transcript. Example wire messages:
{"result":{"speechStarted":{"startTimeMs":0,"confidence":0}}}
{"result":{"speechStopped":{"silenceDurationMs":150}}}speechStopped.silenceDurationMs describes silence when that event was emitted;
it is not a continuously updated idle timer. Use transcription.isFinal to commit
text. For an application idle-stop policy, maintain your own timer and reset it
on speech activity and transcription results. See Turn detection.
Audio capture and runnable examples
Use signed 16-bit little-endian PCM (LINEAR16). Streaming does not accept
MediaRecorder WebM/Opus or Ogg/Opus chunks. Convert captured Float32 samples to
PCM16, and set sampleRateHertz to the actual rate of the samples you send.
16 kHz mono is a convenient starting point; 48 kHz PCM input is also resampled
by the service. Changing the declared rate does not resample the audio.
The samples send 100 ms per chunk: 1,600 samples or 3,200 bytes at 16 kHz mono,
before Base64 encoding. This is an example chunk size, not a protocol limit.
Flush the final shorter chunk before sending closeStream.
For browser capture with an AudioWorklet, keep the processing graph connected to the destination through a zero-gain node, and wait for the worklet to flush its buffered samples before stopping the stream. Release microphone tracks and the audio context on completion or failure.
Start with the Node.js samples or Python samples. Both include file and microphone clients. The sample PCM audio is raw 16 kHz mono PCM16, so you can verify authentication, streaming and shutdown before adding microphone capture.
wss://api.inworld.ai/stt/v1/transcribe:streamBidirectionalBidirectional streaming API for real-time speech-to-text transcription over WebSocket.
This method listens for streaming audio input and returns recognized text chunks one by one as soon as they are ready. Audio chunks are expected to be a part of a single voice input. Suitable for streaming live conversations, microphone input, or other streaming audio sources.
To use the API:
- Send a
transcribeConfigmessage first to configure the session (model, language, audio encoding, etc.). - Stream
audioChunkmessages containing raw audio bytes. - Receive
transcriptionresults as they become available, including both interim (partial) and final results. - Listen for
speechStartedandspeechStoppedevents to detect voice activity changes. - Optionally send
endTurnto signal end of a speaker's turn. - Send
closeStreamwhen done.
Client messages
transcribeConfig
Configure the transcription session. Must be the first message sent. On the wire, wrap the fields below in a top-level transcribeConfig key (see the example) — an unwrapped config is rejected and the socket closes without an error frame.
Payload
modelIdstringrequired
The identifier of the model to use for transcription. Format: "{provider}/{model-name}".
Available models:
inworld/inworld-stt-1— Inworld first-party
See STT Introduction for the full model catalogue.
audioEncodingenum<string>requireddefault: "AUDIO_ENCODING_UNSPECIFIED"
Supported audio encoding formats.
AUDIO_ENCODING_UNSPECIFIED: Not specified. Will return an error.AUTO_DETECT: Automatically detect audio encoding from the audio header.LINEAR16: Uncompressed 16-bit signed little-endian samples (Linear PCM).MP3: MP3 audio. Not supported for streaming transcription.OGG_OPUS: Opus encoded audio wrapped in an OGG container. Not supported for streaming transcription.FLAC: FLAC encoded audio. Lossless format. Not supported for streaming transcription.
Available options:AUDIO_ENCODING_UNSPECIFIEDAUTO_DETECTLINEAR16MP3OGG_OPUSFLAC
languagestring
Language hint in ISO 639 format (e.g., "en", "ja"). Biases the model toward the specified language during automatic language detection. BCP-47 codes (e.g., "en-US") are also accepted and converted to the base language code. The hint additionally constrains the output script for English, Chinese, Cantonese, Japanese, Korean, Russian, and Hindi (e.g. selecting en keeps output in Latin script). See Language Support for the full list of supported languages.
sampleRateHertzinteger
Sample rate of the audio data in Hertz. Required when the sample rate cannot be inferred from the audio header (e.g., raw PCM streams). Default: 16000.
numberOfChannelsinteger
Number of channels in the audio data. Required when the number of channels cannot be inferred from the audio header (e.g., raw PCM streams). Default: 1.
inactivityTimeoutSecondsinteger
Provider-specific inactivity setting. Not applied by inworld/inworld-stt-1; implement application idle-stop timers on the client.
endOfTurnConfidenceThresholdnumber
Confidence threshold for end-of-turn prediction. Higher values reduce false-positives. Range: [0.0, 1.0]. Default: 0.4. See the Turn Detection guide for tuning guidance.
promptsstring[]
Custom vocabulary / key terms. An array of context strings (names, jargon, acronyms) that bias the model toward recognizing these terms. This is a soft bias that helps with ambiguous or uncommon words; it is not a hard keyword lock and does not force exact output. Use letters, digits, spaces, and basic punctuation; other characters (such as #, /, @, or |) are rejected by the gateway with INVALID_ARGUMENT (code 3).
includeWordTimestampsboolean
If true, includes per-word timing information in the response.
enableSpeakerDiarizationboolean
Labels transcribed words with a per-stream speaker identifier. Experimental for inworld/inworld-stt-1 — speaker attribution quality is still improving and speakers may occasionally be misattributed. Set it together with includeWordTimestamps; speaker labels are returned on wordTimestamps[].speaker, and some words may arrive without a label (treat those as unattributed). See the Speaker Diarization guide.
inworldSttV1Configobject
Configuration for Inworld STT 1 models. Set vadThreshold to 0 to disable server-side turn detection and control turn boundaries manually via endTurn — see the Turn Detection guide.
Show child attributes
minEndOfTurnSilenceWhenConfidentinteger
Minimum trailing silence before checking for turn completion, in milliseconds. Applies to streaming. Keep this less than or equal to maxTurnSilence; the minimum is checked before the maximum-silence fallback.
vadThresholdnumber
Voice activity detection threshold for streaming. Range: [0.0, 1.0]. Default: 0.15. Set to 0 for manual turn control.
maxTurnSilenceintegerdefault: 300
Trailing silence that ends the turn even if end-of-turn confidence is insufficient, once minEndOfTurnSilenceWhenConfident has been reached, in milliseconds. Applies to streaming. This is not a transcript-delivery deadline.
voiceProfileConfigobject
Configuration for voice profile detection.
Show child attributes
enableVoiceProfilebooleanrequired
Enables voice profile feature for this request or stream.
topNinteger
Number of top labels from each class to return. Default: 10.
Example
{
"transcribeConfig": {
"modelId": "inworld/inworld-stt-1",
"audioEncoding": "LINEAR16",
"sampleRateHertz": 16000,
"language": "en"
}
}audioChunk
Send a chunk of audio data for transcription. Must be sent after the initial transcribe config message. On the wire, wrap the fields below in a top-level audioChunk key (see the example).
Payload
contentstringrequired
The raw audio bytes in the encoding specified by the transcribe config's audioEncoding.
Example
{
"audioChunk": {
"content": "<BASE64_AUDIO>"
}
}endTurn
Finalize the current turn while keeping the session open for more audio. Send {"endTurn": {}}. At stream shutdown use closeStream instead.
Example
{
"endTurn": {}
}closeStream
Finish audio input and finalize pending audio. Send {"closeStream": {}} once, without an endTurn immediately before it. Keep receiving final transcripts and usage until the server closes the stream.
Example
{
"closeStream": {}
}Server messages
transcription
Transcription result streamed back as audio is processed. May be an interim (partial) result or a final result depending on the isFinal field. Arrives wrapped under a top-level result key as result.transcription (see the example). Each interim replaces the previous interim for the current turn; it is not a delta. Commit a final once and clear the interim.
Payload
transcriptstring
Full transcribed text for this segment.
isFinalboolean
Indicates whether this is a finalized result or an interim (partial) result that may be updated as more audio is processed.
wordTimestampsobject[]
Per-word timing and confidence data. Only populated when includeWordTimestamps is enabled. Not yet supported for inworld/inworld-stt-1.
Show child attributes
wordstring
The transcribed word.
confidencenumber
Recognition confidence score for this word. Range: [0.0, 1.0].
startTimeMsinteger
Offset from the beginning of the audio to the start of this word, in milliseconds.
endTimeMsinteger
Offset from the beginning of the audio to the end of this word, in milliseconds.
speakerinteger
Speaker identifier for this word when diarization is enabled. Integers are assigned per stream in order of first appearance. May be absent on some words even when diarization is enabled — treat unlabeled words as unattributed. See the Speaker Diarization guide.
voiceProfileobjectnull
Voice Profile classification results. null unless voiceProfileConfig.enableVoiceProfile is set in the transcribe config. See the Voice Profiles guide for the category structure.
silenceDurationMsinteger
Duration of trailing silence detected for this result, in milliseconds.
Example
{
"result": {
"transcription": {
"transcript": "Open the door and let me in, please.",
"isFinal": true,
"wordTimestamps": [],
"voiceProfile": null,
"silenceDurationMs": 0
}
}
}usage
Session usage, wrapped as result.usage, delivered during graceful stream completion. Keep reading after closeStream; endTurn alone does not request final session usage. Abrupt disconnection may lose this message.
Payload
transcribedAudioMsinteger
The duration of the transcribed audio in milliseconds.
modelIdstring
The identifier of the model used for transcription.
Example
{
"result": {
"usage": {
"transcribedAudioMs": 2400,
"modelId": "inworld/inworld-stt-1"
}
}
}speechStarted
Signal to indicate the start of a speaker's speech. Sent when voice activity is detected in the audio stream, wrapped under a top-level result key as result.speechStarted.
Payload
startTimeMsinteger
The timestamp of the start of the speech in milliseconds.
confidencenumber
The confidence score of the speech detection.
Example
{
"result": {
"speechStarted": {
"startTimeMs": 0,
"confidence": 0
}
}
}speechStopped
Signal raised when STT detects silence after speech has stopped. Useful for tracking pauses and implementing custom turn-taking logic. Arrives wrapped under a top-level result key as result.speechStopped.
Payload
silenceDurationMsinteger
The duration of silence detected after speech stopped, in milliseconds.
Example
{
"result": {
"speechStopped": {
"silenceDurationMs": 750
}
}
}