STT overview
High-accuracy, low-latency multilingual speech-to-text for live voice agents.
Inworld STT transcribes speech in 30 languages with high accuracy and low latency for live voice agents. Alongside the transcript, it returns Voice Profile signals: age, pitch, emotion, vocal style, and accent. You can tune turn detection or control turns yourself.
Use synchronous transcription for complete audio files, and bidirectional WebSocket streaming for live audio. Both run the inworld/inworld-stt-1 model.
Quickstart
Follow the steps to make your first API call.
STT Playground
Transcribe a file or your microphone in the Portal.
Code examples
Find code examples for common use cases.
Using AI to code? Give your assistant the docs index at https://docs.inworld.ai/llms.txt. For live search, add the MCP server.
Prefer the terminal? Install the Inworld CLI with npm install -g @inworld/cli. Use it to transcribe audio and create API keys. AI agents can use it too.
Your first request
Set INWORLD_API_KEY to the Base64 credential from the Portal, then send a 16 kHz, 16-bit mono WAV file:
import requests
import os
import base64
with open("input.wav", "rb") as f:
audio_content = base64.b64encode(f.read()).decode("utf-8")
response = requests.post(
"https://api.inworld.ai/stt/v1/transcribe",
headers={
"Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}",
"Content-Type": "application/json",
},
json={
"transcribeConfig": {
"modelId": "inworld/inworld-stt-1",
"language": "en",
"audioEncoding": "LINEAR16",
},
"audioData": {"content": audio_content},
},
)
response.raise_for_status()
print(response.json()["transcription"]["transcript"])import fs from "fs";
const audioContent = fs.readFileSync("input.wav").toString("base64");
const response = await fetch("https://api.inworld.ai/stt/v1/transcribe", {
method: "POST",
headers: {
Authorization: `Basic ${process.env.INWORLD_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
transcribeConfig: {
modelId: "inworld/inworld-stt-1",
language: "en",
audioEncoding: "LINEAR16",
},
audioData: { content: audioContent },
}),
});
const result = await response.json();
console.log(result.transcription.transcript);The quickstart explains this request step by step, adds Voice Profile, and shows how to stream with WebSocket.
Explore
Playground guide
Learn how to use the Portal playground.
API reference
Every field on the synchronous and WebSocket endpoints.
Audio formats
Encodings, sample rates, and file-size limits per endpoint.
Language support
The 30 supported languages and how language hints work.
Turn detection
Automatic end-of-turn detection, or client-controlled turns.
Voice profiles
Age, pitch, emotion, vocal style, and accent alongside the transcript.
Speaker diarization
Label who said what. Experimental.
For a spoken conversation, use the Realtime API. It runs STT, an LLM, and TTS in one session. For pricing, see Billing.