Realtime TTS-2 is live. Built for realtime conversation that feels human. Learn more

PvcVoiceService

Train a PVC voice

Starts training on a draft voice's uploaded samples.

Returns the voice resource itself, not a long-running Operation. Poll the Get a PVC voice endpoint to track progress through PVC_VOICE_STATE_QUEUEDPVC_VOICE_STATE_TRAININGPVC_VOICE_STATE_READY (or PVC_VOICE_STATE_FAILED).

Calling Train again while already queued or training is a no-op and returns the voice unchanged.

POST/voices/v1/pvcVoices/{voiceId}:train

This endpoint does not return an Operation and there's no separate operation ID to look up. Poll Get a PVC voice directly.

Requires at least 600 seconds (10 minutes) of cumulative sample audio, measured after any trims are applied.

Training starts are rate-limited per plan; exceeding your plan's rate returns 429.

How long training takes

Wall time scales with how much audio you uploaded, not how many samples it's split across. As a rough guide, training an hour of audio takes on the order of 10-30 minutes; expect it to take longer under heavy platform demand.

Underlying model training is capped at roughly 2800 seconds of audio and 1000 MB of source data per run — audio beyond that is silently truncated rather than rejected. Keep total sample audio comfortably under this if you're uploading long-form recordings.

Authorizations

Authorizationstringrequired

Your API key. Read permissions are required for GET endpoints. Write permissions are required for POST, PATCH, and DELETE endpoints.

For Basic authentication, please populate Basic $INWORLD_API_KEY. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Path Parameters

voiceIdstringrequired

Voice ID of the draft PVC voice to train.

Response

200 - application/json

namestring

Resource name. Format: workspaces/{workspace}/pvcVoices/{voice}.

voiceIdstring

Voice ID, derived from displayName at creation time. Use this value as {voiceId} on every other PVC endpoint, and as the voiceId in TTS synthesis requests once the voice is PVC_VOICE_STATE_READY.

displayNamestring

The human-readable name shown anywhere the voice is listed or selected.

languageCodestring

The voice's language as a BCP-47-shaped locale string, e.g. en-US. Immutable after creation.

stateenum<string>

Lifecycle state of a PVC voice.

  • PVC_VOICE_STATE_DRAFT: Editable. Samples can be added, trimmed, or removed, and metadata can be updated.
  • PVC_VOICE_STATE_QUEUED: Training requested; waiting for a training slot.
  • PVC_VOICE_STATE_TRAINING: Actively training.
  • PVC_VOICE_STATE_READY: Training succeeded. Usable for TTS synthesis. It's permanent and cannot be deleted through this API.
  • PVC_VOICE_STATE_FAILED: Training failed. Editing the voice (e.g. renaming it, or adding/removing a sample) returns it to PVC_VOICE_STATE_DRAFT with its remaining samples intact.

Available options:PVC_VOICE_STATE_UNSPECIFIEDPVC_VOICE_STATE_DRAFTPVC_VOICE_STATE_QUEUEDPVC_VOICE_STATE_TRAININGPVC_VOICE_STATE_READYPVC_VOICE_STATE_FAILED

failureobject

Populated on a PVC voice when its state is PVC_VOICE_STATE_FAILED.

Show child attributes

reasonstring

Machine-readable failure code, e.g. TRAINING_ERROR. New values may be added over time, so don't validate against a hardcoded list. Fall back to displaying message for codes you don't recognize.

messagestring

Human-readable, scrubbed failure message. Never contains uploaded audio, filenames, or transcripts.

incarnationIdstring

Identifier that stays stable across edits to the same voice, and changes each time it is retrained. Use it to tell two reads of the same voiceId apart across a retrain.

samplesobject[]

Audio samples currently attached to the voice.

Show child attributes

sampleIdstring

Sample ID. Use this value as {sampleId} when trimming or deleting the sample.

namestring

Resource name. Format: workspaces/{workspace}/pvcVoices/{voice}/samples/{sample}.

sizeBytesinteger

Size of the uploaded file, in bytes.

durationSecsnumber

Analyzed duration of the sample, in seconds, before any trim is applied.

mimeTypeenum<string>

Detected audio format, sniffed from the file's byte content.

Available options:audio/wavaudio/webmaudio/mpeg

hashstring

Base64-encoded MD5 of the stored object, for verifying upload integrity against the source file.

trimStartMsinteger

Trim start offset in milliseconds, if set.

trimEndMsinteger

Trim end offset in milliseconds, if set.

createTimestring

updateTimestring