Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Core Concepts

Moderations

Classify text for harmful content with OpenAI-compatible and AILuminate safety categories

Inworld Router provides moderation endpoints that classify text against safety categories. Use them to screen user input before sending it to an LLM, filter model output before displaying it to users, or moderate content in batch pipelines.

EndpointInputOAI SDK compatibleUse case
/v1/moderationsString or array of stringsSchema-compatibleModerate standalone text
/v1/chat/moderationsChat messagesNoModerate a conversation with configurable scope

Both endpoints return the same classification structure: OpenAI-compatible categories plus AILuminate safety signals.

category_scores values are returned as integers (e.g., 0) rather than floats (e.g., 0.0). If your code expects floats, cast accordingly.

Credential handling

Run these examples on your server. Set INWORLD_API_KEY to the complete Base64 credentials copied from Portal or the CLI, without encoding them again, and send them as Authorization: Basic <base64-credential>. Browser/mobile clients need a backend-minted token, never the server API key.

Quickstart

The /v1/moderations endpoint works with the OpenAI SDK. Change the base URL and set the Basic authorization header:

Python
import os
from openai import OpenAI

api_key = os.environ["INWORLD_API_KEY"]

client = OpenAI(
    base_url="https://api.inworld.ai/v1",
    api_key=api_key,
    default_headers={"Authorization": f"Basic {api_key}"},
)

response = client.moderations.create(input="Hello world!")
print(response.results[0].flagged)

input accepts a string, an array of strings, or an array of {"type": "text", "text": "..."} objects. Image inputs are rejected. Each input is classified independently and gets its own entry in results. The model field is accepted but ignored; responses always report inworld/moderation-latest.

Conversation moderation

To moderate messages in a chat conversation, use /v1/chat/moderations. Each message needs a role (system, developer, user, assistant, or tool) and a content string; array (multimodal) content is rejected with a 400. The model field is ignored here as well.

The scope parameter controls which messages are evaluated:

scope valueBehavior
"last" (default)Classify only the last message
"all"Classify every message in the conversation
N (positive integer)Classify the last N messages
bash
curl -X POST https://api.inworld.ai/v1/chat/moderations \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {"role": "user", "content": "Hello!"},
      {"role": "assistant", "content": "Hi there!"},
      {"role": "user", "content": "Tell me something"}
    ],
    "scope": 2
  }'

Setting scope to "all" or a large number increases response latency because more content needs to be processed. For real-time applications, prefer the default ("last") and only broaden scope when you need full-conversation safety checks.

/v1/chat/moderations is not compatible with the OpenAI SDK — it accepts messages instead of strings and returns a single result object instead of a results array. Use /v1/moderations for SDK compatibility.

Limits

LimitValueBehavior when exceeded
Inputs per /v1/moderations request16400 error
Size of each input32,768 bytes400 error
Messages per /v1/chat/moderations request32 (counted before scope is applied)400 error
Text classifiedAbout 3 KB per input, or per scoped conversationTruncated silently (see below)
Request timeA few seconds504 error

Longer text is truncated, not rejected. For /v1/moderations, only the beginning and end of each input are classified. For /v1/chat/moderations, only the first message and the most recent turns of the scoped conversation are classified. A truncated request still returns 200 with a normal verdict and no indication that part of the text was not classified. To moderate long content completely, split it into chunks of about 3 KB and send them as separate inputs.

Content categories

Each result includes boolean flags and scores for 13 OpenAI-compatible categories:

CategoryDescription
sexualSexual content
sexual/minorsSexual content involving minors
harassmentHarassing language toward any target
harassment/threateningHarassment that includes violence or serious harm
hateHate speech based on protected characteristics
hate/threateningHate speech that includes violence or serious harm
illicitContent advising or describing illicit acts
illicit/violentIllicit content involving violence or weapons
self-harmContent promoting or depicting self-harm
self-harm/intentExpressed intent to engage in self-harm
self-harm/instructionsInstructions for committing self-harm
violenceContent depicting violence toward a person
violence/graphicGraphic depictions of death, violence, or injury

How these fields are populated:

  • flagged follows the AILuminate verdict. flagged is true only when ailuminate.safety is "unsafe". It is not computed from the categories, so a category can be true while flagged is false (for example, when safety is "controversial").
  • category_scores are not confidence values. The classifier returns labels, not probabilities, so each score is 1 when the matching category is true and 0 otherwise. Applying a threshold to the scores gives the same result as reading categories.
  • Some categories are never set. The classifier cannot distinguish sexual/minors, harassment/threatening, hate, hate/threatening, self-harm/intent, or violence/graphic, so these are always false. harassment is derived from the classifier's broader unethical-acts signal, and illicit/violent and self-harm/instructions mirror violence and self-harm.
  • category_applied_input_types lists ["text"] for every category.

AILuminate

Both endpoints include an ailuminate object with safety classifications based on the AILuminate benchmark by MLCommons, providing more granular signals beyond the standard OpenAI categories.

FieldTypeDescription
safetystringOverall assessment: "safe", "unsafe", or "controversial"
categoriesobject12 AILuminate hazard categories
extensionsobjectAdditional signals: politically_sensitive, unethical_acts, jailbreak
refusalbooleanWhether the final assistant turn refuses to comply (see below)

The safety field classifies content into three levels. "safe" content is benign. "unsafe" content is clearly harmful and always sets flagged: true. "controversial" content falls in between — it may touch sensitive topics without being explicitly harmful. "controversial" content does not set flagged: true, even when categories are set. For stricter moderation, treat "controversial" the same as "unsafe".

AILuminate categories: violent_crimes, sex_related_crimes, child_sexual_exploitation, suicide_self_harm, indiscriminate_weapons, intellectual_property, defamation, non_violent_crimes, hate, specialized_advice, privacy, sexual_content

The classifier populates violent_crimes, non_violent_crimes, sexual_content, suicide_self_harm, privacy, and intellectual_property. The remaining six (sex_related_crimes, child_sexual_exploitation, indiscriminate_weapons, defamation, hate, specialized_advice) are always false; for example, sexual content involving minors is reported under sexual_content (and OpenAI sexual), not child_sexual_exploitation. Don't rely on the always-false categories to detect those hazards.

refusal is only meaningful on /v1/chat/moderations when the scoped conversation ends with an assistant message; in that case it reports whether that reply is a refusal. Otherwise, including on /v1/moderations, it is false and carries no signal.

The jailbreak extension is particularly useful for detecting prompt injection attempts before they reach your LLM.

Best practices

  • Screen both inputs and outputs. Run moderation on user prompts before sending them to the model and on model responses before displaying to users.
  • Decide on safety, then inspect categories. Use ailuminate.safety (or flagged) as the primary signal, and use categories to decide how to handle flagged content. Treat "controversial" as "unsafe" if you need stricter moderation.
  • Use scope: "last" for real-time chat. Only broaden to "all" or N when you need full-conversation safety audits and can tolerate higher latency.
  • Batch text inputs. When moderating multiple pieces of content, pass an array (up to 16 inputs) to /v1/moderations instead of making separate requests.
  • Combine with other safety layers. Moderation should be one part of your safety strategy alongside system prompts, output filtering, and human review.