GPT-OSS-Safeguard 20B available; bring-your-own-policy moderation
GPT-OSS-Safeguard 20B is OpenAI's first open weight reasoning model trained specifically for safety classification tasks, fine-tuned from GPT-OSS, enabling bring-your-own-policy Trust & Safety AI.
Key Features: 131K context window; 65K max output tokens; ~1000 TPS; prompt caching enabled (50% cost savings, $0.037/M vs $0.075/M); Harmony response format with low/medium/high reasoning effort; supports tool use, browser search, code execution, JSON modes, content moderation.
Use Cases: Trust & Safety content moderation, policy-based classification, automated triage, policy testing.
Best Practices: structure policies with Instructions/Definitions/Criteria/Examples sections, keep policies 400-600 tokens, place static content first for caching, use low effort for simple classifications.
Fetched August 11, 2026


