Moderations
Create chat moderation
Classify chat messages for harmful content
/v1/chat/moderationsClassifies messages in a conversation for harmful content. Unlike /v1/moderations, this endpoint accepts chat messages and supports a scope parameter to control which messages are evaluated.
Not compatible with the OpenAI SDK. Use /v1/moderations for SDK compatibility.
Authorizationstringrequired
Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.
Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.
messagesobject[]required
Chat messages to classify, up to 32. Only a few KB of the scoped conversation is classified: the first message and the latest turns. Messages in between are not checked, and the response doesn't say so.
Show child attributes
roleenum<string>required
Available options:systemdeveloperuserassistanttool
contentstringrequired
The message text. Must be a string; content-part arrays are not accepted.
scopeoneOfdefault: "last"
Which messages to classify. "last" (default) processes only the last message. "all" processes every message. A positive integer N processes the last N messages. Be mindful that including many messages increases response latency.
modelstringdefault: "inworld/moderation-latest"
Accepted for compatibility; the Inworld moderation model is always used.
idstring
Unique identifier for the moderation request.
modelstring
The model used for classification.
resultobject
A single aggregated moderation result for the conversation.
Show child attributes
flaggedboolean
true when ailuminate.safety is unsafe. A category can be true while flagged is false, when the content is controversial.
categoriesobject
OpenAI-compatible category flags, derived from the AILuminate classification. sexual/minors, harassment/threatening, hate, hate/threatening, self-harm/intent and violence/graphic are currently always false.
Show child attributes
sexualboolean
sexual/minorsboolean
harassmentboolean
harassment/threateningboolean
hateboolean
hate/threateningboolean
illicitboolean
illicit/violentboolean
self-harmboolean
self-harm/intentboolean
self-harm/instructionsboolean
violenceboolean
violence/graphicboolean
category_scoresobject
1 or 0 for each category, mirroring categories. The classifier returns labels, not calibrated probabilities, so thresholding these scores has no effect. Returned as integers.
Show child attributes
sexualnumber
sexual/minorsnumber
harassmentnumber
harassment/threateningnumber
hatenumber
hate/threateningnumber
illicitnumber
illicit/violentnumber
self-harmnumber
self-harm/intentnumber
self-harm/instructionsnumber
violencenumber
violence/graphicnumber
category_applied_input_typesobject
Input types evaluated for each category.
ailuminateobject
Additional safety classifications based on the AILuminate benchmark by MLCommons. Extends the standard OpenAI moderation categories with more granular signals.
Show child attributes
safetyenum<string>
Overall safety assessment.
Available options:safeunsafecontroversial
categoriesobject
Boolean flags for AILuminate-specific safety categories.
Show child attributes
violent_crimesboolean
sex_related_crimesboolean
child_sexual_exploitationboolean
suicide_self_harmboolean
indiscriminate_weaponsboolean
intellectual_propertyboolean
defamationboolean
non_violent_crimesboolean
hateboolean
specialized_adviceboolean
privacyboolean
sexual_contentboolean
extensionsobject
Additional classification signals.
Show child attributes
politically_sensitiveboolean
unethical_actsboolean
jailbreakboolean
refusalboolean
Whether the content represents a refusal.