inworld/inworld-stt-1) supports 30 languages for speech recognition.
Supported languages
Specifying a language
Thelanguage field is a language hint — it tells the model which language to prefer, but it is not guaranteed to be respected. The model automatically detects the spoken language from the audio, and you can switch languages in the middle of a conversation without changing the hint.
The field accepts ISO 639-1 language codes (e.g., en, ja) matching the codes listed in the table above.
BCP-47 codes (e.g.,
en-US, ja-JP) are also accepted and will be automatically converted to the base ISO 639 language code — for example, en-US becomes en. Regional variants do not affect recognition behavior.For the Inworld first-party model (
inworld/inworld-stt-1), setting language also constrains the output script for English, Chinese, Cantonese, Japanese, Korean, Russian, and Hindi. For example, en keeps the transcript in Latin script (a name spoken in another language is romanized rather than written in its native script), while ja allows Japanese script. This applies only when you set the hint: to let auto-detection follow the spoken language as it changes mid-stream, leave language empty.Third-party provider languages
The Inworld STT API also supports models from third-party providers, each with their own language coverage. See the provider documentation for details:Next steps
Developer Quickstart
Make your first STT API call and get a transcript.
API Reference
View the complete API specification.