- Steering is supported only on
inworld-tts-2. Oninworld-tts-2-flash, steering is not supported — instruction tags and the request-levelinstructionfield are ignored, though non-verbal tags like[laugh]work as usual. For steering, useinworld-tts-2. - Steering instructions must be written in English, even when the input text is in another language.
Free-form turn instructions
Describe the full character of a delivery in natural language, like a director coaching an actor before a take. A single instruction can capture emotion, energy, pacing, and intent all at once. The more fully you describe the delivery — layering mood, physical manner, and intent — the more precisely the voice will perform.[speak as if barely holding back rage forcing every word through gritted teeth] I have told you. Repeatedly. And you STILL didn’t listen.
[overwhelmed with excitement and barely able to contain yourself] We just hit a million users. I still can’t believe it — we actually did it!
[slow and hushed with every word weighted by grief] I got the call this morning. He’s gone.
Metadata-based instructions
Single-property instructions that target one aspect of delivery at a time. The examples below are starting points — feel free to experiment with your own natural language phrasing.-
Articulation: How words are shaped and delivered — add force, crispness, or deliberate rhythm.
Examples:
[say with force][articulate clearly][say with deliberate pauses][articulate clearly] Each step must be followed in order. Do not skip ahead.
-
Intonation: Controls how pitch moves through a phrase — whether it lands decisively or stays open-ended.
Examples:
[say with a falling pitch][say with a rising pitch][say with a falling pitch] That’s my final answer. I’m not changing my mind.
-
Volume: The overall amplitude of the voice, from barely audible to a full, room-filling projection.
Examples:
[very loud][very quiet][very quiet] Don’t make a sound. There’s someone right outside the door.
-
Pitch: The baseline register of the voice — lower for gravity and weight, higher for energy and presence.
Examples:
[say in a low tone][say in a high pitch][say in a high pitch] We just got the green light, the product launches tomorrow!
-
Range: How much pitch varies across the utterance — flat for monotony, expressive for warmth or playfulness.
Examples:
[say playfully][say with no pitch variation][say playfully] So anyway, I was telling her about the trip, and she just laughed the whole time.
-
Speed: The speed of delivery — faster for urgency and tension, slower for weight and clarity.
Examples:
[very fast][very slow][very fast] Run, they’re right behind us, don’t stop, keep moving!
-
Vocal style: Changes how the voice itself sounds — shifting from normal speech into a different mode like whispering or singing.
Examples:
[sing joyfully][whisper in a hushed style][give a nasal quality][sing joyfully] The sun is rising and the world feels new, everything I dreamed is finally coming true.
Non-verbals
Insert organic, human sounds at any point in the text to add realism. Supported tags:[laugh] [breathe] [clear throat] [sigh] [cough] [yawn]
[clear throat] If I could have everyone’s attention, please.
I told him what happened, and he just [laugh] couldn’t believe it!
Emphasis
Capitalize letters within your input text to draw attention to specific words or syllables. Fully capitalizing a word stresses the entire word, while capitalizing individual letters within a word emphasizes a specific syllable.I told you NOT to open that door.
Are you seriously asking if I want pizza? AbsoLUTEly I do.
Common gotchas
I expected the tag to affect only one sentence
I expected the tag to affect only one sentence
An instruction is not scoped to a sentence — it applies from where you write it until you change it. Write
[reset] at the point where normal delivery should resume, or write a new tag there.Instead of: [whisper] Don't move. They're still out there. It's clear now, we can go.Write: [whisper] Don't move. They're still out there. [reset] It's clear now, we can go.I set the instruction field and it stopped applying partway through
I set the instruction field and it stopped applying partway through
An inline tag replaces the
instruction field from the point where the tag appears, and the field does not come back after it. This is the reason to pick one approach rather than combining them.[reset] made my expressive voice sound flat
[reset] made my expressive voice sound flat
It should not.
[reset] removes the instruction you gave, not the character of the voice itself — a voice that is inherently warm, gruff, or excitable keeps that character. If a voice sounds flat after a reset, the expressiveness you were hearing came from the instruction, so use a lighter instruction rather than clearing it entirely.My [tag] was read aloud instead of interpreted
My [tag] was read aloud instead of interpreted
Check the model. Steering is supported only on
inworld-tts-2 — inworld-tts-2-flash ignores instruction tags entirely (non-verbals like [laugh] still work), and older models may read them aloud. Also check the spelling of non-verbals — an unrecognized sound name is treated as an instruction instead of producing a sound.I added a pause and the delivery changed
I added a pause and the delivery changed
It should not, and a
<break/> never clears an instruction. If the delivery changed across a pause, the cause is elsewhere in the text — most often a second [tag] after the break.Best practices
Write instructions in English. Steering instructions must be in English regardless of the language of the input text. For best results, avoid capital letters and punctuation in your instructions. For example,[say quietly in a low tone with deliberate pauses] Bonjour, je suis ravi de vous rencontrer.
Use free-form descriptive instructions for maximum control. The more you describe how you want the voice to perform, the better the output. A bare tag like [sad] gives the model one dimension to work with. A fuller instruction like [say sadly with deliberate pauses in a low voice and hushed style] combines mood, rhythm, pitch, and mode — producing a more nuanced and convincing performance.
Avoid conflicting instructions. Combining opposing directions, for example [whisper in a hushed style] and [very loud] in the same tag, produces unpredictable results. Use one clear instruction per tag.
Match the instruction to the text. The content being spoken should be consistent with the delivery style. A mismatch like [say sadly] applied to This is the happiest day of my life, I just landed my dream job and fell in love! sends contradictory signals and may degrade output quality.
Use one set of instructions per input. Tags that direct delivery, such as articulation, intonation, volume, pitch, range, speed, vocal style, or free-form performance instructions, should appear once at the start. Placing them midway through the text or using multiple instructions throughout will likely produce inconsistent results. Non-verbal tags like [laugh] or [sigh] are the exception and can be inserted inline where the sound should occur.
Avoid: [say in a low tone] I can't believe this happened. [say in a high pitch] Things are looking up though!
Use pause controls for longer pauses. Use pause controls if you want to add longer pauses for added emphasis.