Skip to main content
This guide walks through best practices and techniques for generating high-quality voice clones. For more information on how to create a voice clone, check out this guide. Inworld offers two types of voice cloning: instant voice cloning and professional voice cloning, both available via Inworld Portal. We’ve broken down the best practices in this guide to general best practices that apply to all voice clones, as well as more specific best practices for each type of cloning.

General Best Practices

  1. Capture the full range of expression - Make sure your script and delivery cover the emotions and expressiveness you want the voice to capture. The more variety you include, the better the model will be at recreating those feelings. If the audio is flat, the resulting voice will usually sound monotone as well. Below are some scripts you can use that we’ve found work well:
    • Are you ready to save big? Get set for the sale of the century! Deals and discounts like never before! You won’t want to miss this.
    • Every challenge we face is an opportunity in disguise. Wouldn’t you agree? So cheer up! It’ll all be okay.
    • How have you been? It’s been way too long since we last caught up. By the way, I heard about your recent promotion. Congratulations! I’m so excited for you!
  2. Speak clearly and consistently - Pronounce each word carefully and avoid filler sounds like sighs or coughs. Try not to have unnaturally long pauses in the middle of your recording, as this can affect the flow of the cloned voice.
  3. Minimize noise - Record in a quiet environment and keep a reasonable distance from the microphone to reduce echo, plosives, and device noise. The audio shouldn’t include sound effects or background music. After recording, listen back to ensure your audio is clean and free of any unwanted sounds.

Best Practices for Instant Voice Cloning

  1. Use the full prompt length - Clones work from as little as 3 seconds of audio, but speaker similarity improves with longer samples. Use up to the current 15-second limit where possible (we’re working on supporting longer prompts soon).
  2. Use high-quality audio - Record with at least a 22 kHz sample rate and 16-bit depth.
  3. Vary emotion and delivery - Combine a few short clips that show different expressions into your final clip; use short pauses or crossfades between clips to avoid abrupt cuts.
  4. Use clean audio - Avoid artifacts, background noise, and non-speech sounds.
  5. Normalize volume - Keep levels fairly consistent with normal voice variation; avoid clipping due to very high dB.
  6. Avoid mid-word cuts - Don’t use samples that break in the middle of words.
Instant voice cloning may not perform well for less common voices, such as children’s voices or unique accents. For those use cases, we recommend professional voice cloning.

Best Practices for Professional Voice Cloning

  1. Follow the optimal recording specifications - For the best voice quality, we recommend recording audio with the following specifications:
    • Audio Format: .wav 
    • Sampling Frequency: 48 kHz
    • Bit Rate: 24 bits
    • Codec: Linear PCM (uncompressed)
    • Channel(s): 1 (mono)
    • Loudness Level: -23LUFS ±0.5 LU (compliant with ITU-R BS.1770-3)
    • Peak Values Level (Max): -5 dBFS using True Peak value (compliant with ITU-R BS.1770-3)
    • Noise Floor Level (Max): -60dB
  2. Maintain consistent voice delivery - Keep your voice consistent throughout all recordings. It’s fine to reflect natural variation in speech (such as hesitations, questions, or exclamations), but avoid major changes in accent or style between samples.
  3. Provide ample, high-quality audio - The minimum required audio is 10 minutes total across all uploaded files, but we recommend providing more for the best results. A good mix includes a natural delivery recording plus other recordings (they can be in the same file) that capture different emotions or delivery styles that you would like the final cloned voice to portray—more clean, high-quality audio will generally lead to a higher quality clone.

Automation via API

If you need to clone multiple voices (for example, to support a batch of creators or a pipeline workflow), you can automate voice cloning via the API.
Voice cloning has lower rate limits than regular speech synthesis. For details, see Rate limits.