This is a technical safety and capabilities report for OpenAI's GPT-5 model system, published on arXiv. It is intended for AI researchers, safety evaluators, and developers who need to understand the model's risk profile and mitigation measures.
This document is the official system card for OpenAI's GPT-5 model family, published on arXiv. It is written for AI safety researchers, policy makers, and developers who need a detailed account of the model's safety evaluations, training approach, and risk mitigations. The report covers the full GPT-5 lineup: gpt-5-main, gpt-5-thinking, their mini versions, and the nano and pro variants.
The report begins by introducing the GPT-5 architecture, which combines a fast main model with a deeper reasoning model and a real-time router. It explains the model's training data sources, including public internet data, third-party partnerships, and user-provided information, along with filtering processes to reduce personal data and harmful content. The training methodology for reasoning models uses reinforcement learning to produce internal chains of thought.
A major section covers observed safety challenges and evaluations. It introduces safe-completions, a new safety-training approach that focuses on output safety rather than binary refusals. Evaluations include standard disallowed content tests, production benchmarks with multi-turn conversations, and assessments of sycophancy, jailbreaks, instruction hierarchy, prompt injections, hallucinations, deception, image input safety, health-related outputs, multilingual performance, and fairness via…
The report details red teaming and external assessments, including expert red teaming for violent attack planning and prompt injections. It then presents the Preparedness Framework, with capability assessments for biological and chemical risks (using benchmarks like ProtocolQA and TroubleshootingBench, plus external evaluations by SecureBio), cybersecurity (Capture the Flag challenges, cyber range tests, and evaluations by Pattern Labs), and AI self-improvement (SWE-bench, MLE-Bench,…
A dedicated section addresses safeguards for high biological and chemical risk, including a threat model and biological threat taxonomy, safeguard design across model training, system-level protections, account-level enforcement, API access controls, and a Trusted Access Program. Testing covers model safety training, system-level protections, expert red teaming for bioweaponization, third-party red teaming, and external government red teaming. The report concludes with security controls and an…