Core Concepts
Moderations
Classify text for harmful content with OpenAI-compatible and AILuminate safety categories
Inworld Router provides moderation endpoints that classify text against safety categories. Use them to screen user input before sending it to an LLM, filter model output before displaying it to users, or moderate content in batch pipelines.
| Endpoint | Input | OAI SDK compatible | Use case |
|---|---|---|---|
/v1/moderations | String or array of strings | Schema-compatible | Moderate standalone text |
/v1/chat/moderations | Chat messages | No | Moderate a conversation with configurable scope |
Both endpoints return the same classification structure: OpenAI-compatible categories plus AILuminate safety signals.
category_scores values are returned as integers (e.g., 0) rather than floats (e.g., 0.0). If your code expects floats, cast accordingly.
Credential handling
Run these examples on your server. Set INWORLD_API_KEY to the complete Base64 credentials copied from Portal or the CLI, without encoding them again, and send them as Authorization: Basic <base64-credential>. Browser/mobile clients need a backend-minted token, never the server API key.
Quickstart
The /v1/moderations endpoint works with the OpenAI SDK. Change the base URL and set the Basic authorization header:
import os
from openai import OpenAI
api_key = os.environ["INWORLD_API_KEY"]
client = OpenAI(
base_url="https://api.inworld.ai/v1",
api_key=api_key,
default_headers={"Authorization": f"Basic {api_key}"},
)
response = client.moderations.create(input="Hello world!")
print(response.results[0].flagged)import OpenAI from 'openai';
const apiKey = process.env.INWORLD_API_KEY;
const client = new OpenAI({
baseURL: 'https://api.inworld.ai/v1',
apiKey,
defaultHeaders: { Authorization: `Basic ${apiKey}` },
});
const response = await client.moderations.create({ input: 'Hello world!' });
console.log(response.results[0].flagged);curl -X POST https://api.inworld.ai/v1/moderations \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": "Hello world!"}'input accepts a string, an array of strings, or an array of {"type": "text", "text": "..."} objects. Image inputs are rejected. Each input is classified independently and gets its own entry in results. The model field is accepted but ignored; responses always report inworld/moderation-latest.
Conversation moderation
To moderate messages in a chat conversation, use /v1/chat/moderations. Each message needs a role (system, developer, user, assistant, or tool) and a content string; array (multimodal) content is rejected with a 400. The model field is ignored here as well.
The scope parameter controls which messages are evaluated:
scope value | Behavior |
|---|---|
"last" (default) | Classify only the last message |
"all" | Classify every message in the conversation |
N (positive integer) | Classify the last N messages |
curl -X POST https://api.inworld.ai/v1/chat/moderations \
-H "Authorization: Basic $INWORLD_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hi there!"},
{"role": "user", "content": "Tell me something"}
],
"scope": 2
}'Setting scope to "all" or a large number increases response latency because more content needs to be processed. For real-time applications, prefer the default ("last") and only broaden scope when you need full-conversation safety checks.
/v1/chat/moderations is not compatible with the OpenAI SDK — it accepts messages instead of strings and returns a single result object instead of a results array. Use /v1/moderations for SDK compatibility.
Limits
| Limit | Value | Behavior when exceeded |
|---|---|---|
Inputs per /v1/moderations request | 16 | 400 error |
| Size of each input | 32,768 bytes | 400 error |
Messages per /v1/chat/moderations request | 32 (counted before scope is applied) | 400 error |
| Text classified | About 3 KB per input, or per scoped conversation | Truncated silently (see below) |
| Request time | A few seconds | 504 error |
Longer text is truncated, not rejected. For /v1/moderations, only the beginning and end of each input are classified. For /v1/chat/moderations, only the first message and the most recent turns of the scoped conversation are classified. A truncated request still returns 200 with a normal verdict and no indication that part of the text was not classified. To moderate long content completely, split it into chunks of about 3 KB and send them as separate inputs.
Content categories
Each result includes boolean flags and scores for 13 OpenAI-compatible categories:
| Category | Description |
|---|---|
sexual | Sexual content |
sexual/minors | Sexual content involving minors |
harassment | Harassing language toward any target |
harassment/threatening | Harassment that includes violence or serious harm |
hate | Hate speech based on protected characteristics |
hate/threatening | Hate speech that includes violence or serious harm |
illicit | Content advising or describing illicit acts |
illicit/violent | Illicit content involving violence or weapons |
self-harm | Content promoting or depicting self-harm |
self-harm/intent | Expressed intent to engage in self-harm |
self-harm/instructions | Instructions for committing self-harm |
violence | Content depicting violence toward a person |
violence/graphic | Graphic depictions of death, violence, or injury |
How these fields are populated:
flaggedfollows the AILuminate verdict.flaggedistrueonly whenailuminate.safetyis"unsafe". It is not computed from the categories, so a category can betruewhileflaggedisfalse(for example, whensafetyis"controversial").category_scoresare not confidence values. The classifier returns labels, not probabilities, so each score is1when the matching category istrueand0otherwise. Applying a threshold to the scores gives the same result as readingcategories.- Some categories are never set. The classifier cannot distinguish
sexual/minors,harassment/threatening,hate,hate/threatening,self-harm/intent, orviolence/graphic, so these are alwaysfalse.harassmentis derived from the classifier's broader unethical-acts signal, andillicit/violentandself-harm/instructionsmirrorviolenceandself-harm. category_applied_input_typeslists["text"]for every category.
AILuminate
Both endpoints include an ailuminate object with safety classifications based on the AILuminate benchmark by MLCommons, providing more granular signals beyond the standard OpenAI categories.
| Field | Type | Description |
|---|---|---|
safety | string | Overall assessment: "safe", "unsafe", or "controversial" |
categories | object | 12 AILuminate hazard categories |
extensions | object | Additional signals: politically_sensitive, unethical_acts, jailbreak |
refusal | boolean | Whether the final assistant turn refuses to comply (see below) |
The safety field classifies content into three levels. "safe" content is benign. "unsafe" content is clearly harmful and always sets flagged: true. "controversial" content falls in between — it may touch sensitive topics without being explicitly harmful. "controversial" content does not set flagged: true, even when categories are set. For stricter moderation, treat "controversial" the same as "unsafe".
AILuminate categories: violent_crimes, sex_related_crimes, child_sexual_exploitation, suicide_self_harm, indiscriminate_weapons, intellectual_property, defamation, non_violent_crimes, hate, specialized_advice, privacy, sexual_content
The classifier populates violent_crimes, non_violent_crimes, sexual_content, suicide_self_harm, privacy, and intellectual_property. The remaining six (sex_related_crimes, child_sexual_exploitation, indiscriminate_weapons, defamation, hate, specialized_advice) are always false; for example, sexual content involving minors is reported under sexual_content (and OpenAI sexual), not child_sexual_exploitation. Don't rely on the always-false categories to detect those hazards.
refusal is only meaningful on /v1/chat/moderations when the scoped conversation ends with an assistant message; in that case it reports whether that reply is a refusal. Otherwise, including on /v1/moderations, it is false and carries no signal.
The jailbreak extension is particularly useful for detecting prompt injection attempts before they reach your LLM.
Best practices
- Screen both inputs and outputs. Run moderation on user prompts before sending them to the model and on model responses before displaying to users.
- Decide on
safety, then inspect categories. Useailuminate.safety(orflagged) as the primary signal, and usecategoriesto decide how to handle flagged content. Treat"controversial"as"unsafe"if you need stricter moderation. - Use
scope: "last"for real-time chat. Only broaden to"all"orNwhen you need full-conversation safety audits and can tolerate higher latency. - Batch text inputs. When moderating multiple pieces of content, pass an array (up to 16 inputs) to
/v1/moderationsinstead of making separate requests. - Combine with other safety layers. Moderation should be one part of your safety strategy alongside system prompts, output filtering, and human review.