Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Moderations

Create moderation

Classify text inputs for harmful content

POST/v1/moderations

Classifies one or more text inputs for harmful content. Schema-compatible with the OpenAI Moderations API — works with the OpenAI SDK.

The response includes standard OpenAI moderation categories as well as AILuminate safety classifications for more granular signals.

Authorizations

Authorizationstringrequired

Your authentication credentials. For Basic authentication, please populate Basic $INWORLD_API_KEY.

Please make sure your API Key has write permissions for the Router API in order to create, update, and delete routers. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Body

application/json

inputoneOfrequired

A string, an array of strings, or an array of {"type": "text", "text": ...} parts to classify. Image inputs are not supported. Up to 16 inputs of up to 32 KB each. Only about the first and last few KB of a long input are classified; the rest is not checked, and the response doesn't say so.

modelstringdefault: "inworld/moderation-latest"

Accepted for OpenAI SDK compatibility; the Inworld moderation model is always used.

Response

200 - application/json

idstring

Unique identifier for the moderation request.

modelstring

The model used for classification.

resultsobject[]

Array of moderation results, one per input.

Show child attributes

flaggedboolean

true when ailuminate.safety is unsafe. A category can be true while flagged is false, when the content is controversial.

categoriesobject

OpenAI-compatible category flags, derived from the AILuminate classification. sexual/minors, harassment/threatening, hate, hate/threatening, self-harm/intent and violence/graphic are currently always false.

Show child attributes

sexualboolean

sexual/minorsboolean

harassmentboolean

harassment/threateningboolean

hateboolean

hate/threateningboolean

illicitboolean

illicit/violentboolean

self-harmboolean

self-harm/intentboolean

self-harm/instructionsboolean

violenceboolean

violence/graphicboolean

category_scoresobject

1 or 0 for each category, mirroring categories. The classifier returns labels, not calibrated probabilities, so thresholding these scores has no effect. Returned as integers.

Show child attributes

sexualnumber

sexual/minorsnumber

harassmentnumber

harassment/threateningnumber

hatenumber

hate/threateningnumber

illicitnumber

illicit/violentnumber

self-harmnumber

self-harm/intentnumber

self-harm/instructionsnumber

violencenumber

violence/graphicnumber

category_applied_input_typesobject

Input types evaluated for each category.

ailuminateobject

Additional safety classifications based on the AILuminate benchmark by MLCommons. Extends the standard OpenAI moderation categories with more granular signals.

Show child attributes

safetyenum<string>

Overall safety assessment.

Available options:safeunsafecontroversial

categoriesobject

Boolean flags for AILuminate-specific safety categories.

Show child attributes

violent_crimesboolean

sex_related_crimesboolean

child_sexual_exploitationboolean

suicide_self_harmboolean

indiscriminate_weaponsboolean

intellectual_propertyboolean

defamationboolean

non_violent_crimesboolean

hateboolean

specialized_adviceboolean

privacyboolean

sexual_contentboolean

extensionsobject

Additional classification signals.

Show child attributes

politically_sensitiveboolean

unethical_actsboolean

jailbreakboolean

refusalboolean

Whether the content represents a refusal.