> ## Documentation Index
> Fetch the complete documentation index at: https://docs.inworld.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Train a PVC voice

> Starts training on a draft voice's uploaded samples.
> 
> Returns the voice resource itself, **not** a long-running Operation. Poll the Get a PVC voice endpoint to track progress through `PVC_VOICE_STATE_QUEUED` → `PVC_VOICE_STATE_TRAINING` → `PVC_VOICE_STATE_READY` (or `PVC_VOICE_STATE_FAILED`).
> 
> Calling Train again while already queued or training is a no-op and returns the voice unchanged.

<Warning>
**This endpoint does not return an Operation** and there's no separate operation ID to look up. Poll [Get a PVC voice](https://docs.inworld.ai/api-reference/pvcAPI/pvcvoiceservice/get-pvc-voice.md) directly.
</Warning>

Requires at least **600 seconds (10 minutes)** of cumulative sample audio, measured after any [trims](https://docs.inworld.ai/api-reference/pvcAPI/pvcvoiceservice/trim-pvc-voice-sample.md) are applied.

Training starts are rate-limited per plan; exceeding your plan's rate returns `429`, and the error message names the limit. Wait before trying again rather than retrying in a loop. See [Plan limits](https://docs.inworld.ai/tts/professional-voice-cloning.md#plan-limits) for each plan's allowance.

A voice that is already `PVC_VOICE_STATE_READY` can't be trained again and returns `409`. To retrain, [create a new PVC voice](https://docs.inworld.ai/api-reference/pvcAPI/pvcvoiceservice/create-pvc-voice.md) — see [Retrain or remove a trained voice](https://docs.inworld.ai/tts/professional-voice-cloning.md#retrain-or-remove-a-trained-voice).

## How long training takes

Wall time scales with how much audio you uploaded, not how many samples it's split across. As a rough guide, training an hour of audio takes on the order of 10-30 minutes; expect it to take longer under heavy platform demand.

<Note>
Underlying model training is capped at roughly 2800 seconds of audio and 1000 MB of source data per run — audio beyond that is silently truncated rather than rejected. Keep total sample audio comfortably under this if you're uploading long-form recordings.
</Note>

## API reference

**Endpoint:** `POST https://api.inworld.ai/voices/v1/pvcVoices/{voiceId}:train`

### Authorization

- `Authorization` (string; required) — Your [API key](../../../api-reference/introduction). Read permissions are required for GET endpoints. Write permissions are required for POST, PATCH, and DELETE endpoints.
  
   For Basic authentication, please populate `Basic $INWORLD_API_KEY`. You can create a key in one command with the [Inworld CLI](../../../developer-tools/inworld-cli): `inworld workspace add-key`.

### Path parameters

- `voiceId` (string; required) — Voice ID of the draft PVC voice to train.

### Request body

Content type: `application/json`


#### Request example

```json
{}
```

### Response

Status: `200`. Content type: `application/json`.

- `name` (string) — Resource name. Format: `workspaces/{workspace}/pvcVoices/{voice}`.
- `voiceId` (string) — Voice ID, derived from `displayName` at creation time. Use this value as `{voiceId}` on every other PVC endpoint, and as the `voiceId` in TTS synthesis requests once the voice is `PVC_VOICE_STATE_READY`.
- `displayName` (string) — The human-readable name shown anywhere the voice is listed or selected.
- `languageCode` (string) — The voice's language as a BCP-47-shaped locale string, e.g. `en-US`. Immutable after creation.
- `state` (enum<string>; options: "PVC_VOICE_STATE_UNSPECIFIED", "PVC_VOICE_STATE_DRAFT", "PVC_VOICE_STATE_QUEUED", "PVC_VOICE_STATE_TRAINING", "PVC_VOICE_STATE_READY", "PVC_VOICE_STATE_FAILED") — Lifecycle state of a PVC voice.
  
  - `PVC_VOICE_STATE_DRAFT`: Editable. Samples can be added, trimmed, or removed, and metadata can be updated.
  - `PVC_VOICE_STATE_QUEUED`: Training requested; waiting for a training slot.
  - `PVC_VOICE_STATE_TRAINING`: Actively training.
  - `PVC_VOICE_STATE_READY`: Training succeeded. Usable for TTS synthesis. It's permanent and cannot be deleted through this API.
  - `PVC_VOICE_STATE_FAILED`: Training failed. Editing the voice (e.g. renaming it, or adding/removing a sample) returns it to `PVC_VOICE_STATE_DRAFT` with its remaining samples intact.
- `failure` (object) — Populated on a PVC voice when its `state` is `PVC_VOICE_STATE_FAILED`.
  - `reason` (string) — Machine-readable failure code, e.g. `TRAINING_ERROR`. New values may be added over time, so don't validate against a hardcoded list. Fall back to displaying `message` for codes you don't recognize.
  - `message` (string) — Human-readable, scrubbed failure message. Never contains uploaded audio, filenames, or transcripts.
- `incarnationId` (string) — Identifier that stays stable across edits to the same voice, and changes each time it is retrained. Use it to tell two reads of the same `voiceId` apart across a retrain.
- `samples` (object[]) — Audio samples currently attached to the voice.
  - `sampleId` (string) — Sample ID. Use this value as `{sampleId}` when trimming or deleting the sample.
  - `name` (string) — Resource name. Format: `workspaces/{workspace}/pvcVoices/{voice}/samples/{sample}`.
  - `sizeBytes` (integer) — Size of the uploaded file, in bytes.
  - `durationSecs` (number) — Analyzed duration of the sample, in seconds, before any trim is applied.
  - `mimeType` (enum<string>; options: "audio/wav", "audio/webm", "audio/mpeg") — Detected audio format, sniffed from the file's byte content.
  - `hash` (string) — Base64-encoded MD5 of the stored object, for verifying upload integrity against the source file.
  - `trimStartMs` (integer) — Trim start offset in milliseconds, if set.
  - `trimEndMs` (integer) — Trim end offset in milliseconds, if set.
- `createTime` (string)
- `updateTime` (string)

### Response examples

#### 200: queued

```json
{
  "name": "workspaces/your_workspace_id/pvcVoices/my-professional-voice",
  "voiceId": "my-professional-voice",
  "displayName": "my-professional-voice",
  "languageCode": "en-US",
  "state": "PVC_VOICE_STATE_QUEUED",
  "incarnationId": "a1b2c3d4",
  "createTime": "2026-08-31T12:00:00Z",
  "updateTime": "2026-08-31T12:10:00Z"
}
```

#### 400: insufficient_audio

```json
{
  "code": 3,
  "message": "invalid request: at least 600 seconds of sample audio is required to train, got 214",
  "details": []
}
```

#### 429: default_response

```json
{
  "code": 0,
  "message": "string",
  "details": [
    {
      "@type": "string"
    }
  ]
}
```

#### default: default_response

```json
{
  "code": 0,
  "message": "string",
  "details": [
    {
      "@type": "string"
    }
  ]
}
```

### Code examples

#### cURL

```bash
curl --location --request POST 'https://api.inworld.ai/voices/v1/pvcVoices/<voice-id>:train' \
--header "Authorization: Basic $INWORLD_API_KEY" \
--header 'Content-Type: application/json' \
--data '{}'
```

#### Python

```python
import requests

voice_id = "<voice-id>"
url = f"https://api.inworld.ai/voices/v1/pvcVoices/{voice_id}:train"
headers = {
    "Authorization": "Basic <api-key>",
    "Content-Type": "application/json"
}

response = requests.post(url, headers=headers, json={})
print(response.json())
```

#### JavaScript

```javascript
const voiceId = '<voice-id>';
const url = `https://api.inworld.ai/voices/v1/pvcVoices/${voiceId}:train`;

const response = await fetch(url, {
  method: 'POST',
  headers: {
    'Authorization': 'Basic <api-key>',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({}),
});

const data = await response.json();
console.log(data);
```
