VoiceService
Clone a voice
Clone a voice from audio samples.
/voices/v1/voices:cloneShort URL Path: /workspaces/{workspace} is no longer required in the path for simplicity and clarity. When omitted, the workspace is derived from your API key. The previous URL with the full path /voices/v1/workspaces/{workspace}/voices:clone would continue to be supported.
Setting gender, ageGroup, and categories. These cannot be set at clone time — sending them here (or putting them in tags) is silently ignored, and they come back empty on the response. After cloning, set them with UpdateVoice.
Choosing the voice's language
Set the language with languageCode — the canonical locale string, e.g. "en-US", "en-GB", "vi":
- Matching is forgiving — case- and separator-insensitive (
en-gb,EN_GB, anden-GBare equivalent). A bare language code with no region ("en","pt") selects the language's default accent. - Accent is part of the locale — there is no separate accent field. To clone a British-accented voice, send
"en-GB"; for Australian,"en-AU". - Auto-detect — omit the field (or send
"auto") to detect the language from the audio samples. - Validation — values outside the supported catalog are rejected with
INVALID_ARGUMENT; nothing is silently coerced.
See Languages for the supported languages.
Legacy langCode. Older integrations set the language via the langCode enum (the locale with - replaced by _, uppercased — en-GB → EN_GB; AUTO = auto-detect). It remains accepted, and responses populate it alongside languageCode. Set at most one of the two on a request; use languageCode in new code.
<RequestField body="languageCode" type="string">
Canonical locale string, e.g. en-US, en-GB, vi. See Choosing the voice's language above.
</RequestField>
Authorizations
Authorizationstringrequired
Your API key. Read permissions are required for GET endpoints. Write permissions are required for POST, PATCH, and DELETE endpoints.
For Basic authentication, please populate Basic $INWORLD_API_KEY. You can create a key in one command with the Inworld CLI: inworld workspace add-key.
Body
application/jsondisplayNamestringrequired
The human-readable name shown anywhere the voice is listed or selected. Keep it short and distinctive so users can find it easily.
langCodeenum<string>
Legacy enum encoding of the voice's language. The full accepted set is much larger than the values listed here: every supported locale has an enum name (the locale with - replaced by _, uppercased — en-GB becomes EN_GB). Prefer the languageCode string field on new integrations. AUTO (or omitting the language entirely) auto-detects the language.
Available options:EN_USZH_CNKO_KRJA_JPRU_RUAUTOIT_ITES_ESPT_BRDE_DEFR_FRAR_SAPL_PLNL_NLHI_INHE_IL
languageCodestring
The voice's language as a canonical BCP-47-shaped locale string (e.g. en-US, en-GB, vi). Set at most one of languageCode or langCode — they are two encodings of the same value. Matching is case- and separator-insensitive (en-gb, EN_GB and en-GB are equivalent); a bare language code with no region (e.g. en, pt) selects the language's default accent. Omit both fields to auto-detect the language (equivalently: langCode: "AUTO" or languageCode: "auto"). Values outside the supported catalog are rejected with INVALID_ARGUMENT. See Languages for the supported set.
voiceSamplesobject[]required
Voice samples used for cloning. For best results, provide clear audio and avoid speaking in multiple languages, whispering, or making non-verbal sounds like coughing. Instant voice cloning works best with a 10-30 sec audio clip; longer clips will be cut off at 30 sec, which can affect quality. See Voice Cloning Best Practices for guidance on how to generate a high-quality voice clone.
Show child attributes
audioDatastringrequired
Binary audio data for the sample (base64-encoded in JSON). Supports WAV and MP3 formats.
transcriptionstring
Optional user-provided transcription of the audio sample. If one is not provided, the transcription will be generated automatically.
descriptionstring
Longer blurb that explains the voice's tone, accent, use cases, or other relevant attributes. Helpful for search and selection.
tagsstring[]
Free-form labels for filtering, grouping, and discovery (e.g. ["british", "calm"]). This is not where gender or age go — those are separate fields set via UpdateVoice after cloning.
audioProcessingConfigobject
Audio processing config for voice cloning.
Show child attributes
removeBackgroundNoiseboolean
Whether to remove background noise from the samples. If true, an audio isolation model will be used to clean the samples. Note: This can degrade quality if samples are already clean.