Get Started
Asynchronous Transcription
Submit a long recording as a job and collect the transcript when it finishes
Synchronous transcription answers in one request, which works while a recording is small enough to sit in a request body and short enough that you are willing to wait. Asynchronous transcription is for everything else: you hand over a recording, receive a job, and collect the transcript when the job finishes.
Use it for recordings measured in minutes or hours — meetings, interviews, podcasts, call archives. For live audio, use streaming transcription instead; for a short clip, the synchronous endpoint is simpler.
How it works
Sending the audio
Three ways, and the right one depends mostly on size.
| Way | Field | Use it when |
|---|---|---|
| Inline | audioData.content | The recording is small. Audio is base64-encoded, which makes the request about a third larger than the file. |
| Multipart | multipart/form-data | The recording is large. The file is streamed rather than held whole. |
| By URL | audioUri | The audio already lives somewhere reachable, such as cloud storage. |
Inline
curl -X POST https://api.inworld.ai/stt/v1/transcribe:async \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcribeConfig": {
"modelId": "inworld/inworld-stt-1",
"audioEncoding": "LINEAR16"
},
"audioData": { "content": "<base64-encoded audio>" }
}'Multipart
Send the config first and the file second. The upload is streamed as it arrives, so a config that comes after the audio is read too late.
curl -X POST https://api.inworld.ai/stt/v1/transcribe:async \
-H "Authorization: Basic $INWORLD_API_KEY" \
-F 'transcribeConfig={"modelId":"inworld/inworld-stt-1","audioEncoding":"LINEAR16"};type=application/json' \
-F "file=@recording.wav"By URL
{
"transcribeConfig": { "modelId": "inworld/inworld-stt-1", "audioEncoding": "LINEAR16" },
"audioUri": "https://storage.googleapis.com/your-bucket/recording.wav?..."
}The URL must be https, must serve the audio directly, and must be reachable without your credentials — either public, or carrying its own authorization such as a signed cloud-storage URL. Redirects are refused, so give the URL that serves the bytes rather than one that points at it. The service fetches the audio while your submit request is still in flight, so a URL it cannot reach is reported as a failed submit rather than a failed job — and submitting a long recording this way takes as long as the fetch does.
Most convenient sharing links redirect — file-sharing services, shortened URLs, and console URLs such as storage.cloud.google.com. Use the direct serving URL, for example https://storage.googleapis.com/<bucket>/<object>.
Configuration
transcribeConfig takes the same fields as synchronous transcription, with two differences.
Every audio encoding is accepted, including the compressed formats streaming rejects — MP3, FLAC, OGG_OPUS — because the audio is a stored file rather than a live stream. AUTO_DETECT works where the file carries headers.
Live-session fields are ignored: inactivityTimeoutSeconds and endOfTurnConfidenceThreshold mean nothing to a recording that is already complete.
Speaker diarization, voice profiles and word timestamps all work as they do synchronously. Their output lands in the transcript document rather than in the operation — see below.
Reading the result
While the job runs, the operation reports done: false. When it finishes:
{
"name": "workspaces/my-workspace/sttTranscriptionJobs/6f1c.../operations/1790...",
"done": true,
"response": {
"resultUri": "https://storage.googleapis.com/...?X-Goog-Signature=...",
"expireTime": "2026-10-01T12:00:00Z",
"language": "English",
"usage": { "transcribedAudioMs": 7200000 }
}
}A job that failed carries an error instead of a response, with a code and a message. Either way the operation is done, so polling ends.
resultUri points at the transcript document:
{
"jobId": "6f1c...",
"transcript": "Full text of the recording...",
"language": "English",
"segments": [
{ "endTimeMs": "7200000", "transcript": "Full text of the recording..." }
],
"usage": { "transcribedAudioMs": 7200000, "modelId": "inworld/inworld-stt-1" }
}Word timestamps and speaker labels
Asking for includeWordTimestamps adds a wordTimestamps array to each segment, and enableSpeakerDiarization adds a speaker to each word in it:
{
"segments": [
{
"endTimeMs": "25861",
"transcript": "Furthermore, our universal fit phone cases...",
"wordTimestamps": [
{ "word": "Furthermore", "startTimeMs": 60, "endTimeMs": 870, "speaker": 0 },
{ "word": "our", "startTimeMs": 1240, "endTimeMs": 1420, "speaker": 0 }
]
}
]
}Speaker identifiers are small integers scoped to the one job: they mark "same speaker in this recording", not a person you can recognise across jobs.
Voice profiles arrive the same way — asking for voiceProfileConfig.enableVoiceProfile adds a voiceProfile object to each segment, with the same categories and { label, confidence } shape as a synchronous response.
language is the detected language's name, not a BCP-47 tag, and the millisecond fields are 64-bit integers, which JSON carries as strings. A field that is zero — such as the first segment's startTimeMs — is omitted rather than sent.
segments currently contains a single entry spanning the whole recording. The field is a list so that finer segmentation can be added without changing the document's shape — do not rely on there being exactly one entry.
Limits
| Limit | Value |
|---|---|
| Audio per job | 512 MB |
| Jobs in flight per account | Depends on your plan |
| Result link validity | 24 hours from completion |
| Result retention | 7 days from completion, after which the transcript is deleted |
| Submitted audio retention | 7 days, after which the recording is deleted |
audioUri fetch timeout | 5 minutes, and the submit waits for it |
Submitting while you are already at your in-flight limit is refused with RESOURCE_EXHAUSTED, before the audio is transferred. Wait for a job to finish and submit again.
Download the transcript within 24 hours. The link is signed once, when the job completes, and polling the operation again returns the same link rather than a fresh one — so once it expires the transcript cannot be retrieved, even though it is kept for seven days. Keep your own copy if you need it later; neither the transcript nor the submitted audio is a permanent store.
Errors
| What you see | What it means |
|---|---|
401 / 403 | The API key is missing, wrong, or belongs to a different environment. |
RESOURCE_EXHAUSTED | Either the in-flight job limit, or the request rate limit. The message says which. |
INVALID_ARGUMENT | A malformed request — check audioEncoding, modelId, and the audioUri rules above. |
FAILED_PRECONDITION | The audioUri could not be fetched. The request was well formed; the fetch was not. It is answered on the submit, since that is when the fetch happens. |