TextToSpeech
Synthesize speech (batch)
Submit many independent synthesis requests as one background job and immediately receive a long-running operation. Poll the operation via the Get operation endpoint until done is true, then download the results file named by resultsUri — it lists every item's outcome, keyed by the customId you assigned. Each item's request is an ordinary synthesis request, validated exactly as the async endpoint would validate it; one invalid item rejects the whole batch at submit. That includes timestampType in a language without alignment support — like async jobs, a batch item never returns audio with requested timestamps silently missing.
/tts/v1/voice:synthesizeBatchResults file
When the batch is done, the operation's response.resultsUri names this JSON file. By default it is the file itself, and each item's artifacts are pre-signed download URLs. For results delivered to your storage with outputConfig.resultsUri, or hosted with outputConfig.packaging: ZIP, it is a zip archive holding the file at response.resultsPath, and each item's artifacts are paths inside that archive.
totalItemsinteger
Number of items submitted.
completedItemsinteger
Number of items that produced audio.
failedItemsinteger
Number of items that failed. Omitted rather than sent as 0 when every item succeeded, so read it with a default.
resultsobject[]
One entry per submitted item, in submission order. Correlate entries by customId, not by position.
Show child attributes — results
customIdstring
The key you assigned this item on submit.
audioUristring
Pre-signed download URL for this item's audio, in the format its request asked for, when the results are packaged as files. Fetch it without an Authorization header, before the response's expireTime. Absent when the item failed.
timestampsUristring
Pre-signed download URL for this item's timestamp alignment JSON, when the results are packaged as files. Present only when the item's request set timestampType and the item succeeded. See Reading the timestamps document for its form.
audioPathstring
Path of this item's audio inside the archive, for example items/0.mp3, when the results are packaged as a zip. Absent when the item failed.
timestampsPathstring
Path of this item's timestamp alignment JSON inside the archive, for example items/0.timestamps.json, when the results are packaged as a zip. Present only when the item's request set timestampType and the item succeeded. See Reading the timestamps document for its form.
errorobject
Why this item produced no audio: a status with a numeric code and a message. Absent when the item succeeded.
Authorizationstringrequired
Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY. You can create a key in one command with the Inworld CLI: inworld workspace add-key.
itemsobject[]required
The requests to synthesize, between 1 and 10000. The whole request must also fit the 16 MiB message limit, which caps the aggregate around 4M characters — the item ceiling is only reachable when items are short.
Show child attributes
customIdstringrequired
Your key for this item, echoed into the results file. Required, and unique within the batch — it is the only thing correlating a result back to what you submitted. Treat it as opaque; the service never interprets it.
requestoneOfrequired
The synthesis request for this item, identical in shape to a synchronous request and subject to the same validation. On-Demand accounts are additionally capped at 10,000 characters across the whole batch.
Show child attributes
textstringrequired
The text to be synthesized into speech. Maximum input of 100,000 characters.
voiceIdstring
The ID of the voice to use for synthesizing speech. Set either voiceId or voiceDesign, not both.
audioConfigobject
Configurations to use when synthesizing speech.
Show child attributes
audioEncodingenum<string>default: "MP3"
The desired output format of the synthesized audio. Defaults to MP3.
LINEAR16: Uncompressed 16-bit signed little-endian samples (Linear PCM). For non-streaming, the WAV header is included in the response. For streaming, the WAV header is included in every audio chunk.MP3: MP3 audio.OGG_OPUS: Opus encoded audio wrapped in an ogg container. The result will be a file which can be played natively on Android, and in browsers (at least Chrome and Firefox). The quality of the encoding is considerably higher than MP3 while using approximately the same bitrate.ALAW: ALAW encoded audio. 8-bit companded PCM.MULAW: MULAW encoded audio. 8-bit companded PCM.FLAC: FLAC encoded audio. Lossless audio format.PCM: PCM audio. Uncompressed 16-bit signed little-endian samples with no WAV header.WAV: WAV audio. Uncompressed 16-bit signed little-endian samples. For non-streaming, the WAV header is included in the response. For streaming, the WAV header is included in the first audio chunk only.
Available options:LINEAR16MP3OGG_OPUSALAWMULAWFLACPCMWAV
bitRateinteger
Bits per second of the audio. Only for compressed audio formats (MP3, OGG_OPUS). The default is 128,000.
sampleRateHertzinteger
The synthesis sample rate (in hertz) for this audio. Accepts values within the range [8000, 48000]. Supported sample rates are: 8000, 16000, 22050, 24000, 32000, 44100, 48000.
When this is specified, if this is different from the voice's natural sample rate, then the audio will be converted to the desired sample rate (which might result in worse audio quality), unless the specified sample rate is not supported for the encoding chosen, in which case it will fail the request and return an error. The default is 48,000.
speakingRatenumber
Speaking rate/speed, in the range [0.5, 1.5]. The default is 1.0, which is the normal native speed supported by the specific voice. We recommend using values above 0.8 to ensure high quality.
modelIdstringrequired
The ID of the model to use for synthesizing speech. See Models for available models.
languagestring
BCP-47 language tag (e.g., en-US, fr-FR, ja-JP) specifying the language that the given voice should speak the text in. Matching is case- and separator-insensitive for standard two-part tags (en-gb, EN_GB, and en-GB are equivalent); longer tags with extension subtags must match a catalog entry exactly. If a localized voice prompt exists for the language, it will be used. When omitted, the original voice prompt will be used and the language will be auto-detected from the input text. If an invalid language code is provided, an error will be returned.
See Languages for more details.
deliveryModeenum<string>default: "DELIVERY_MODE_UNSPECIFIED"
Only supported by `inworld-tts-2`. The field is ignored on other models.
Controls how varied the output is.
DELIVERY_MODE_UNSPECIFIED: Defaults toBALANCEDbehavior.STABLE: Optimizes for more consistent, predictable output.BALANCED: Balanced between stability and diversity.CREATIVE: Optimizes for increased emotional range and variation.
Available options:DELIVERY_MODE_UNSPECIFIEDSTABLEBALANCEDCREATIVE
instructionstring
Only supported by `inworld-tts-2`. The field is ignored on other models.
Speaking-style instruction for this request — for example speak loudly and urgently or sound out of breath. Applies to the whole request. Write it in English, even when text is in another language. An empty string means unset.
You can also change the instruction mid-text with inline [bracket] tags. A tag applies from where it appears until you change it, so it overrides this field from that point on; [reset] removes the instruction for the rest of the text. Prefer one approach or the other rather than combining them.
See Steering for the full guide.
temperaturenumberdefault: 1
Ignored on `inworld-tts-2`. Use [`deliveryMode`](#body-delivery-mode) instead.
Determines the degree of randomness when sampling audio tokens to generate the response.
Defaults to 1.0. Accepts values between 0 (exclusive) and 2 (inclusive). Higher values will make the output more random and can lead to more expressive results. Lower values will make it more deterministic. If 0 is provided, the default value will be used.
For the most stable results, we recommend using the default value.
timestampTypeenum<string>default: "TIMESTAMP_TYPE_UNSPECIFIED"
Controls timestamp metadata returned with the audio. When enabled, the response includes timing arrays, which can be useful for word-highlighting, karaoke-style captions, and lipsync.
- WORD: Output arrays under
timestampInfo.wordAlignment(words, wordStartTimeSeconds, wordEndTimeSeconds). - CHARACTER: Output arrays under
timestampInfo.characterAlignment(characters, characterStartTimeSeconds, characterEndTimeSeconds). - TIMESTAMPTYPEUNSPECIFIED: Do not compute alignment; timestamp arrays will be empty or omitted.
Phonetic details: phoneticDetails is currently only returned for WORD alignment (not CHARACTER).
Latency note: Alignment adds additional computation. Enabling alignment can increase latency.
Available options:TIMESTAMP_TYPE_UNSPECIFIEDWORDCHARACTER
applyTextNormalizationenum<string>default: "APPLY_TEXT_NORMALIZATION_UNSPECIFIED"
When enabled, text normalization automatically expands and standardizes things like numbers, dates, times, and abbreviations before converting them to speech. For example, Dr. Smith becomes Doctor Smith, and 3/10/25 is spoken as March tenth, twenty twenty-five. Turning this off may reduce latency, but the speech output will read the text exactly as written. Defaults to automatically deciding whether to apply text normalization.
Available options:APPLY_TEXT_NORMALIZATION_UNSPECIFIEDONOFF
enhanceGenerationbooleandefault: false
When true, applies denoising to the synthesized audio to reduce background noise and artifacts, improving the overall audio quality of the generation. Defaults to false (no denoising).
synthesisContextobject
Context for the current synthesis request. Supplying the text of earlier requests from the same session or conversation gives the model additional context and can improve the quality of the generation, especially for short or ambiguous input text. Context text is not billed. The texts of all previous requests combined must not exceed 2,000 characters.
Show child attributes
previousRequestsobject[]
Previous requests from the same session or conversation, in the order they were synthesized.
Show child attributes
textstring
The text that was synthesized in the previous request.
voiceDesignobject
Preview. Design a voice for this synthesis request without creating or saving a voice. Supported by inworld-tts-2 on unary, server-streaming, async, and batch synthesis. Omit voiceId. Each request or batch item designs its own voice; repeating a description does not guarantee the same voice. See Ad-hoc voice design.
Show child attributes
designPromptstringrequired
A description of the desired voice, such as its accent, pitch, timbre, and delivery. Use 7–1,024 characters; leading and trailing whitespace does not count toward the minimum. This describes the voice; put the words to speak in text.
outputConfigobject
Preview. Where the whole batch's results go and how they are packaged. Set it here, at the top level: an item whose request sets outputConfig is rejected. Omit it to have Inworld host the results file and per-item download URLs. See Results delivery.
Show child attributes
resultsUristring
A URL to an object in storage you own, pre-signed for an HTTP PUT: an Amazon S3 presigned URL, a Google Cloud Storage V4 signed URL, or an Azure SAS URL that grants write. The job uploads its results there as one zip archive, and nothing is retained by Inworld. The signature must stay valid for at least 12 hours after submit for an async job, and 48 hours for a batch. Omit it to have Inworld host the results for about 7 days. Required when the workspace has Zero Data Retention enabled: a job without it is refused with FAILED_PRECONDITION.
packagingenum<string>default: "PACKAGING_UNSPECIFIED"
How the results are packaged.
- ZIP: One zip archive holding
results.jsonbeside every artifact it names, by paths relative to it. The response names the archive inresultsUriand the document inresultsPath. - FILES: Each artifact as its own pre-signed download URL. Only for results Inworld hosts:
FILEStogether withresultsUriis rejected withINVALID_ARGUMENT. - PACKAGING_UNSPECIFIED:
ZIPwhenresultsUriis set,FILESotherwise.
Available options:PACKAGING_UNSPECIFIEDZIPFILES
namestring
Server-assigned operation resource name, in the format workspaces/{workspace}/ttsBatchJobs/{batch}/operations/{operation}. Pass it verbatim as the path of the Get operation endpoint to poll for completion.
metadataobject
SynthesizeSpeechBatchMetadata counting the items finished so far, present from the moment the batch is accepted through to the terminal response.
Show child attributes
@typestring
totalItemsinteger
Number of items in the batch. Fixed when the batch is accepted, so it is readable before any work starts.
completedItemsinteger
Items synthesized successfully so far. A zero counter is spelled out as 0 on the submit response but omitted from the poll response, so read it with a default rather than testing whether the key is present.
failedItemsinteger
Items that finished with an error so far. Item errors do not fail the batch — it still completes, and every item keeps its own entry in the results file. Omitted from the poll response while zero.
createTimestring
When the batch was accepted.
doneboolean
If false, the job is still running. If true, the job has finished and exactly one of error or response is set.
errorobject
Show child attributes
codeinteger
The error code, as specified by gRPC status codes.
messagestring
A short description of the error.
detailsobject[]
Show child attributes
@typestring
responseobject
Set when the batch succeeded: a SynthesizeSpeechBatchResponse naming the results file.
Show child attributes
@typestring
Type of the serialized message, always type.googleapis.com/ai.inworld.tts.v1.SynthesizeSpeechBatchResponse.
resultsUristring
Where the results live. Packaged as FILES, the default for results Inworld hosts: a pre-signed download URL for the results file. Packaged as ZIP: the archive holding that file at resultsPath — your outputConfig.resultsUri without its signature when the results were delivered to your storage, or a pre-signed download URL for an archive Inworld hosts.
resultsPathstring
Path of the results file inside the archive at resultsUri. Absent when the results are packaged as FILES, where resultsUri is the results file itself.
expireTimestring
Time at which the download URLs expire — resultsUri and, packaged as FILES, every audio URL inside the results file — approximately 7 days after the batch completes. Absent when the results were delivered to outputConfig.resultsUri.