Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Get Started

Asynchronous Transcription

Submit a long recording as a job and collect the transcript when it finishes

Synchronous transcription answers in one request, which works while a recording is small enough to sit in a request body and short enough that you are willing to wait. Asynchronous transcription is for everything else: you hand over a recording, receive a job, and collect the transcript when the job finishes.

Use it for recordings measured in minutes or hours — meetings, interviews, podcasts, call archives. For live audio, use streaming transcription instead; for a short clip, the synchronous endpoint is simpler.

How it works

Submit the recording

POST /stt/v1/transcribe:async returns immediately with an operation — a handle naming the job, not the transcript.

Poll the operation

GET /lro/v1alpha/{operation name} reports whether the job has finished. Poll every few seconds; a long recording takes a while.

Download the transcript

A finished operation carries a link to the transcript. The link is signed and needs no credentials.

Sending the audio

Three ways, and the right one depends mostly on size.

WayFieldUse it when
InlineaudioData.contentThe recording is small. Audio is base64-encoded, which makes the request about a third larger than the file.
Multipartmultipart/form-dataThe recording is large. The file is streamed rather than held whole.
By URLaudioUriThe audio already lives somewhere reachable, such as cloud storage.

Inline

bash
curl -X POST https://api.inworld.ai/stt/v1/transcribe:async \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcribeConfig": {
      "modelId": "inworld/inworld-stt-1",
      "audioEncoding": "LINEAR16"
    },
    "audioData": { "content": "<base64-encoded audio>" }
  }'

Multipart

Send the config first and the file second. The upload is streamed as it arrives, so a config that comes after the audio is read too late.

bash
curl -X POST https://api.inworld.ai/stt/v1/transcribe:async \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -F 'transcribeConfig={"modelId":"inworld/inworld-stt-1","audioEncoding":"LINEAR16"};type=application/json' \
  -F "file=@recording.wav"

By URL

json
{
  "transcribeConfig": { "modelId": "inworld/inworld-stt-1", "audioEncoding": "LINEAR16" },
  "audioUri": "https://storage.googleapis.com/your-bucket/recording.wav?..."
}

The URL must be https, must serve the audio directly, and must be reachable without your credentials — either public, or carrying its own authorization such as a signed cloud-storage URL. Redirects are refused, so give the URL that serves the bytes rather than one that points at it. The service fetches the audio while your submit request is still in flight, so a URL it cannot reach is reported as a failed submit rather than a failed job — and submitting a long recording this way takes as long as the fetch does.

Most convenient sharing links redirect — file-sharing services, shortened URLs, and console URLs such as storage.cloud.google.com. Use the direct serving URL, for example https://storage.googleapis.com/<bucket>/<object>.

Configuration

transcribeConfig takes the same fields as synchronous transcription, with two differences.

Every audio encoding is accepted, including the compressed formats streaming rejects — MP3, FLAC, OGG_OPUS — because the audio is a stored file rather than a live stream. AUTO_DETECT works where the file carries headers.

Live-session fields are ignored: inactivityTimeoutSeconds and endOfTurnConfidenceThreshold mean nothing to a recording that is already complete.

Speaker diarization, voice profiles and word timestamps all work as they do synchronously. Their output lands in the transcript document rather than in the operation — see below.

Reading the result

While the job runs, the operation reports done: false. When it finishes:

json
{
  "name": "workspaces/my-workspace/sttTranscriptionJobs/6f1c.../operations/1790...",
  "done": true,
  "response": {
    "resultUri": "https://storage.googleapis.com/...?X-Goog-Signature=...",
    "expireTime": "2026-10-01T12:00:00Z",
    "language": "English",
    "usage": { "transcribedAudioMs": 7200000 }
  }
}

A job that failed carries an error instead of a response, with a code and a message. Either way the operation is done, so polling ends.

resultUri points at the transcript document:

json
{
  "jobId": "6f1c...",
  "transcript": "Full text of the recording...",
  "language": "English",
  "segments": [
    { "endTimeMs": "7200000", "transcript": "Full text of the recording..." }
  ],
  "usage": { "transcribedAudioMs": 7200000, "modelId": "inworld/inworld-stt-1" }
}

Word timestamps and speaker labels

Asking for includeWordTimestamps adds a wordTimestamps array to each segment, and enableSpeakerDiarization adds a speaker to each word in it:

json
{
  "segments": [
    {
      "endTimeMs": "25861",
      "transcript": "Furthermore, our universal fit phone cases...",
      "wordTimestamps": [
        { "word": "Furthermore", "startTimeMs": 60, "endTimeMs": 870, "speaker": 0 },
        { "word": "our", "startTimeMs": 1240, "endTimeMs": 1420, "speaker": 0 }
      ]
    }
  ]
}

Speaker identifiers are small integers scoped to the one job: they mark "same speaker in this recording", not a person you can recognise across jobs.

Voice profiles arrive the same way — asking for voiceProfileConfig.enableVoiceProfile adds a voiceProfile object to each segment, with the same categories and { label, confidence } shape as a synchronous response.

language is the detected language's name, not a BCP-47 tag, and the millisecond fields are 64-bit integers, which JSON carries as strings. A field that is zero — such as the first segment's startTimeMs — is omitted rather than sent.

segments currently contains a single entry spanning the whole recording. The field is a list so that finer segmentation can be added without changing the document's shape — do not rely on there being exactly one entry.

Limits

LimitValue
Audio per job512 MB
Jobs in flight per accountDepends on your plan
Result link validity24 hours from completion
Result retention7 days from completion, after which the transcript is deleted
Submitted audio retention7 days, after which the recording is deleted
audioUri fetch timeout5 minutes, and the submit waits for it

Submitting while you are already at your in-flight limit is refused with RESOURCE_EXHAUSTED, before the audio is transferred. Wait for a job to finish and submit again.

Download the transcript within 24 hours. The link is signed once, when the job completes, and polling the operation again returns the same link rather than a fresh one — so once it expires the transcript cannot be retrieved, even though it is kept for seven days. Keep your own copy if you need it later; neither the transcript nor the submitted audio is a permanent store.

Errors

What you seeWhat it means
401 / 403The API key is missing, wrong, or belongs to a different environment.
RESOURCE_EXHAUSTEDEither the in-flight job limit, or the request rate limit. The message says which.
INVALID_ARGUMENTA malformed request — check audioEncoding, modelId, and the audioUri rules above.
FAILED_PRECONDITIONThe audioUri could not be fetched. The request was well formed; the fetch was not. It is answered on the submit, since that is when the fetch happens.