Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

STT overview

High-accuracy, low-latency multilingual speech-to-text for live voice agents.

Inworld STT transcribes speech in 30 languages with high accuracy and low latency for live voice agents. Alongside the transcript, it returns Voice Profile signals: age, pitch, emotion, vocal style, and accent. You can tune turn detection or control turns yourself.

Use synchronous transcription for complete audio files, and bidirectional WebSocket streaming for live audio. Both run the inworld/inworld-stt-1 model.

Using AI to code? Give your assistant the docs index at https://docs.inworld.ai/llms.txt. For live search, add the MCP server.

Prefer the terminal? Install the Inworld CLI with npm install -g @inworld/cli. Use it to transcribe audio and create API keys. AI agents can use it too.

Your first request

Set INWORLD_API_KEY to the Base64 credential from the Portal, then send a 16 kHz, 16-bit mono WAV file:

Python
import requests
import os
import base64

with open("input.wav", "rb") as f:
    audio_content = base64.b64encode(f.read()).decode("utf-8")

response = requests.post(
    "https://api.inworld.ai/stt/v1/transcribe",
    headers={
        "Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}",
        "Content-Type": "application/json",
    },
    json={
        "transcribeConfig": {
            "modelId": "inworld/inworld-stt-1",
            "language": "en",
            "audioEncoding": "LINEAR16",
        },
        "audioData": {"content": audio_content},
    },
)
response.raise_for_status()
print(response.json()["transcription"]["transcript"])

The quickstart explains this request step by step, adds Voice Profile, and shows how to stream with WebSocket.

Explore

For a spoken conversation, use the Realtime API. It runs STT, an LLM, and TTS in one session. For pricing, see Billing.