Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

STT quickstart

Make your first Inworld STT API request

In this quickstart, you'll send an audio file to the STT API and receive a transcript. It also highlights Inworld STT (inworld/inworld-stt-1), which adds Voice Profile (age, pitch, emotion, vocal style, accent) and configurable turn-taking (automatic or manual).

Make your first STT API request

Create an API key

Create an Inworld account.

In Inworld Portal, generate an API key by going to Settings > API Keys. Copy the Base64 credentials.

Set your API key as an environment variable.

macOS and Linux
export INWORLD_API_KEY='your-base64-api-key-here'

Prepare an audio file

The STT API accepts audio in several formats (e.g. MP3, OGG_OPUS, FLAC, LINEAR16). Audio bytes are sent in the request payload as a base64-encoded string — base64 is the transport encoding, not the audio format. Requirements vary by use case:

Use caseFormatNotes
File upload (sync)LINEAR16, MP3, OGG_OPUS, FLAC, AUTO_DETECTSample rate can be auto-detected from file headers when possible
StreamingLINEAR16 (PCM)Other encodings are not supported for streaming to minimize latency and preserve quality

Recommended settings:

  • Sample rate: 16,000 Hz (STT performs best at this rate; lower sample rates like 8 kHz contain fewer data points, reducing accuracy)
  • Bit depth: 16-bit (for LINEAR16)
  • Channels: Mono (1 channel)

For file uploads (MP3, FLAC, OGG_OPUS, WAV), sampleRateHertz is optional — the API can auto-detect it from the file header.

Sync transcription accepts audio files up to ~16 MB. The actual duration depends on the encoding (e.g., ~18 minutes of MP3 or ~8 minutes of 16 kHz 16-bit WAV). For larger files, split them into chunks or use the WebSocket streaming endpoint.

Send the request

Audio is sent as a JSON payload with base64-encoded audio content. The API returns the complete transcript when processing is complete (and optionally Voice Profile, when returned by the API).

Create a new file inworld_stt_quickstart.py or inworld_stt_quickstart.js and use the code below. The Inworld model (inworld/inworld-stt-1) provides transcription plus optional Voice Profile (age, pitch, emotion, vocal style, accent) and configurable turn-taking for streaming.

Python
import requests
import os
import base64

# Sync endpoint
URL = "https://api.inworld.ai/stt/v1/transcribe"

# Use a 16-bit PCM WAV file (16 kHz, mono)
with open("input.wav", "rb") as f:
    audio_content = base64.b64encode(f.read()).decode("utf-8")

payload = {
    "transcribeConfig": {
        "modelId": "inworld/inworld-stt-1",
        "language": "en",
        "audioEncoding": "LINEAR16",
        "voiceProfileConfig": {
            "enableVoiceProfile": True,
        },
    },
    "audioData": {"content": audio_content},
}

headers = {
    "Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}",
    "Content-Type": "application/json",
}

response = requests.post(URL, headers=headers, json=payload)
response.raise_for_status()
result = response.json()

print("Transcript:", result["transcription"]["transcript"])

# Voice Profile (when returned by the API)
if "voiceProfile" in result and result["voiceProfile"]:
    vp = result["voiceProfile"]
    if vp.get("age"):
        print("Age:", vp["age"].get("label"), vp["age"].get("confidence"))
    if vp.get("pitch"):
        print("Pitch:", vp["pitch"].get("label"), vp["pitch"].get("confidence"))

Review the response

The response includes the transcript and usage fields, plus optional voiceProfile when available.

Response (sync)

FieldDescription
transcription.transcriptThe transcribed text
transcription.isFinalWhether the result is finalized
transcription.wordTimestampsPer-word timing data (when available)
usageUsage metrics for billing
voiceProfile(When returned) Age, pitch, emotion, vocalStyle, accent with label and confidence

Configuration parameters

transcribeConfig

FieldTypeRequiredDescription
modelIdstringYesSTT model ID. Use inworld/inworld-stt-1 for WebSocket and HTTP
languagestringNoISO 639 language code (e.g. en, ja). BCP-47 codes like en-US are also accepted and converted to the base language. If omitted, the model may auto-detect. See Language Support for the full list
audioEncodingstringYesOne of: LINEAR16, MP3, OGG_OPUS, FLAC, AUTO_DETECT. For streaming, use LINEAR16 only
sampleRateHertzintegerNoSample rate in Hz. Default 16000. Can be omitted for formats with headers (MP3, FLAC, OGG_OPUS, WAV)
numberOfChannelsintegerNoChannel count. Default 1
promptsstring[]NoCustom vocabulary / key terms (names, jargon, acronyms) that bias recognition toward those terms. A soft bias, not a hard keyword lock. Max 100 terms per request; letters, digits, spaces, and . , ' ( ) : ! ? ; - only — a term with any other character rejects the whole request. Supported across models. The Realtime API equivalent is audio.input.transcription.prompts
voiceProfileConfigobjectNoVoice Profile configuration. See below

voiceProfileConfig

FieldTypeRequiredDescription
enableVoiceProfileboolYesSet to true to enable Voice Profile analysis
topNintegerNoNumber of top labels per category to return. Default: 10

audioData

FieldTypeRequiredDescription
contentstringYesBase64-encoded audio bytes

Run the code

Python
pip install requests  # if needed
python inworld_stt_quickstart.py

Example output:

Transcript: Hey, I just wanted to check in on the delivery status for my order.

Streaming (WebSocket)

For real-time microphone or live audio, follow the WebSocket integration guide for authentication, browser capture, message examples, and runnable Node.js/Python samples:

  1. First message must contain transcribeConfig (same fields as above, including voiceProfileConfig to enable Voice Profile).
  2. Later messages send audioChunk with base64-encoded LINEAR16 (PCM) audio only.
  3. Turn and stream end:
    • To finalize a speaker turn and continue sending audio, send endTurn.
    • When done, flush the last audio chunk and send closeStream once, without a preceding endTurn. Keep reading final transcripts and usage until the server closes the stream.

Example first WebSocket message:

json
{
  "transcribeConfig": {
    "modelId": "inworld/inworld-stt-1",
    "audioEncoding": "LINEAR16"
  }
}

Interim transcripts replace the previous interim for the current turn; they are not deltas. Append each final once and clear the interim.

Responses stream back as Transcription (interim and final), optional voiceProfile, speech events (speechStarted when voice activity is detected, speechStopped when silence is detected after speech), and finally Usage when the stream is closed. Every server message arrives wrapped in a result envelope (errors arrive as {"error": {"code": 3, "message": "invalid transcribe config"}}):

json
{
  "result": {
    "transcription": {
      "transcript": "Open the door and let me in, please.",
      "isFinal": true,
      "wordTimestamps": [],
      "voiceProfile": null,
      "silenceDurationMs": 0
    }
  }
}

Streaming endpoint (WebSocket): wss://api.inworld.ai/stt/v1/transcribe:streamBidirectional

Next Steps