Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Capabilities

Language tags

Speak each part of a text in the language you choose with inline lang tags.

A single text often mixes languages: a language tutor quoting a Spanish word, a travel guide naming a Japanese temple, an assistant reading a French title. Wrap each part of the text in <lang> tags that name its language, and each part is spoken in that language, with that language's pronunciation.

<lang lang="en-US">In Spanish, the word</lang> <lang lang="es-ES">llama</lang> <lang lang="en-US">means flame.</lang>

"In Spanish, the word" and "means flame" are spoken in English, and "llama" in Spanish.

Language tags work with inworld-tts-2 on every endpoint: non-streaming, streaming, WebSocket, async and batch.

When to use language tags

Use language tags when you need a specific language spoken in the middle of a sentence, and especially for text that is written the same way in two languages. Without a tag, the model has to guess which reading you mean.

  • Chinese and Japanese share characters. 日本大学 is Rìběn dàxué in Chinese and Nihon daigaku in Japanese:
    <lang lang="zh-CN">这所学校的日语名字是</lang><lang lang="ja-JP">日本大学。</lang>
  • English and Spanish share many spellings. llama is the animal in English, but in Spanish it means "flame" and is pronounced differently, as in the example above.

Syntax

  • Open a span with <lang lang="CODE"> and close it with </lang>, where CODE is a BCP-47 language tag, such as es, es-MX or ja-JP. See supported languages.
  • The SSML spelling <lang xml:lang="CODE"> is accepted too, so SSML generated by other tools works unchanged.
  • Tag and attribute names are case insensitive, and the code may use single, double or no quotes.
  • Tags do not nest. Close each span before opening the next.
  • A </lang> with no open span, and a tag whose code is not a valid language tag, are ignored.

WebSocket

How you send language tags on a WebSocket context depends on how the context buffers text:

  • autoMode: false: make each flush self-contained. Wrap every part of the flushed text in the language it should be spoken in, and close every span you open before you flush.
  • autoMode: true with CLIENT_SEGMENTED: make each sendText payload self-contained in the same way.
  • autoMode: true with SENTENCE_BOUNDARY: send text as it arrives, including individual tokens from an LLM. A tag may be split across messages, and a span may continue across several.

Tags count toward the characters billed for a request, as other markup does.