Realtime TTS-2 is live. Built for realtime conversation that feels human. Learn more

PvcVoiceService

Create a PVC voice

Creates a new Professional Voice Clone in PVC_VOICE_STATE_DRAFT. Upload audio samples, then start training. See the steps below.

POST/voices/v1/pvcVoices

Language. Only en-US is currently supported for Professional Voice Cloning. languageCode is immutable after creation.

A draft voice has no audio yet. Build it out with the rest of this API before training:

Create the draft voice

This endpoint. Returns a voiceId derived from displayName.

Upload audio samples

Upload PVC voice samples — repeat until you have at least 10 minutes of cumulative audio. Optionally trim individual samples with Trim a PVC voice sample.

Train

Train a PVC voice starts training and moves the voice to PVC_VOICE_STATE_QUEUED.

Poll until ready

Get a PVC voice to watch state progress to PVC_VOICE_STATE_READY (or PVC_VOICE_STATE_FAILED). Once ready, use the voiceId anywhere you'd use a regular voice, e.g. Synthesize speech.

Each voice claims one of your plan's PVC voice slots for as long as it exists. Check your usage with Resolve upload limits (usedPvcVoiceSlots / maxPvcVoiceSlots); when all slots are taken, creation is refused until a slot is freed or the plan is upgraded. On-Demand accounts also need a payment method on file.

For recording and preparation tips, see Voice Cloning best practices.

Authorizations

Authorizationstringrequired

Your API key. Read permissions are required for GET endpoints. Write permissions are required for POST, PATCH, and DELETE endpoints.

For Basic authentication, please populate Basic $INWORLD_API_KEY. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Body

application/json

displayNamestringrequired

The human-readable name shown anywhere the voice is listed or selected. The voice's voiceId is derived from this value; renaming the voice later does not change its voiceId.

languageCodestring

The voice's language as a BCP-47-shaped locale string. Currently only en-US is supported for Professional Voice Cloning; defaults to en-US if omitted. Immutable after creation.

Response

200 - application/json

namestring

Resource name. Format: workspaces/{workspace}/pvcVoices/{voice}.

voiceIdstring

Voice ID, derived from displayName at creation time. Use this value as {voiceId} on every other PVC endpoint, and as the voiceId in TTS synthesis requests once the voice is PVC_VOICE_STATE_READY.

displayNamestring

The human-readable name shown anywhere the voice is listed or selected.

languageCodestring

The voice's language as a BCP-47-shaped locale string, e.g. en-US. Immutable after creation.

stateenum<string>

Lifecycle state of a PVC voice.

  • PVC_VOICE_STATE_DRAFT: Editable. Samples can be added, trimmed, or removed, and metadata can be updated.
  • PVC_VOICE_STATE_QUEUED: Training requested; waiting for a training slot.
  • PVC_VOICE_STATE_TRAINING: Actively training.
  • PVC_VOICE_STATE_READY: Training succeeded. Usable for TTS synthesis. It's permanent and cannot be deleted through this API.
  • PVC_VOICE_STATE_FAILED: Training failed. Editing the voice (e.g. renaming it, or adding/removing a sample) returns it to PVC_VOICE_STATE_DRAFT with its remaining samples intact.

Available options:PVC_VOICE_STATE_UNSPECIFIEDPVC_VOICE_STATE_DRAFTPVC_VOICE_STATE_QUEUEDPVC_VOICE_STATE_TRAININGPVC_VOICE_STATE_READYPVC_VOICE_STATE_FAILED

failureobject

Populated on a PVC voice when its state is PVC_VOICE_STATE_FAILED.

Show child attributes

reasonstring

Machine-readable failure code, e.g. TRAINING_ERROR. New values may be added over time, so don't validate against a hardcoded list. Fall back to displaying message for codes you don't recognize.

messagestring

Human-readable, scrubbed failure message. Never contains uploaded audio, filenames, or transcripts.

incarnationIdstring

Identifier that stays stable across edits to the same voice, and changes each time it is retrained. Use it to tell two reads of the same voiceId apart across a retrain.

samplesobject[]

Audio samples currently attached to the voice.

Show child attributes

sampleIdstring

Sample ID. Use this value as {sampleId} when trimming or deleting the sample.

namestring

Resource name. Format: workspaces/{workspace}/pvcVoices/{voice}/samples/{sample}.

sizeBytesinteger

Size of the uploaded file, in bytes.

durationSecsnumber

Analyzed duration of the sample, in seconds, before any trim is applied.

mimeTypeenum<string>

Detected audio format, sniffed from the file's byte content.

Available options:audio/wavaudio/webmaudio/mpeg

hashstring

Base64-encoded MD5 of the stored object, for verifying upload integrity against the source file.

trimStartMsinteger

Trim start offset in milliseconds, if set.

trimEndMsinteger

Trim end offset in milliseconds, if set.

createTimestring

updateTimestring