Both endpoints return the same classification structure: OpenAI-compatible categories plus AILuminate safety signals.
category_scores values are returned as integers (e.g., 0) rather than floats (e.g., 0.0). If your code expects floats, cast accordingly.Quickstart
The/v1/moderations endpoint works directly with the OpenAI SDK — just change the base URL:
Conversation moderation
To moderate messages in a chat conversation, use/v1/chat/moderations. The scope parameter controls which messages are evaluated:
/v1/chat/moderations is not compatible with the OpenAI SDK — it accepts messages instead of strings and returns a single result object instead of a results array. Use /v1/moderations for SDK compatibility.Content categories
Each result includes boolean flags and numeric confidence scores for 13 categories:
The
flagged field is true when any category exceeds the default threshold. Use category_scores (0–1 confidence values) to set custom thresholds for your application.
AILuminate
Both endpoints include anailuminate object with safety classifications based on the AILuminate benchmark by MLCommons, providing more granular signals beyond the standard OpenAI categories.
The
safety field classifies content into three levels. "safe" content is benign. "unsafe" content is clearly harmful and always sets flagged: true. "controversial" content falls in between — it may touch sensitive topics without being explicitly harmful. By default, "controversial" content is treated as safe and does not set flagged: true. For stricter moderation, treat "controversial" the same as "unsafe".
AILuminate categories: violent_crimes, sex_related_crimes, child_sexual_exploitation, suicide_self_harm, indiscriminate_weapons, intellectual_property, defamation, non_violent_crimes, hate, specialized_advice, privacy, sexual_content
The jailbreak extension is particularly useful for detecting prompt injection attempts before they reach your LLM.
Best practices
- Screen both inputs and outputs. Run moderation on user prompts before sending them to the model and on model responses before displaying to users.
- Use
category_scoresfor custom thresholds. Theflaggedboolean uses default thresholds. For your application, you may want stricter thresholds for certain categories (e.g.,sexual/minors) and more permissive ones for others. - Use
scope: "last"for real-time chat. Only broaden to"all"orNwhen you need full-conversation safety audits and can tolerate higher latency. - Batch text inputs. When moderating multiple pieces of content, pass an array to
/v1/moderationsinstead of making separate requests. - Combine with other safety layers. Moderation should be one part of your safety strategy alongside system prompts, output filtering, and human review.