OpenAI: Safety justification should be submitted before cutting-edge reinforcement learning training
On September 28, 2026, OpenAI published a safety-related article, stating that before continuing any cutting-edge reinforcement learning training, structured safety documentation should be required. Ideally, such documentation should reach the evidence-based structured risk argument level used in safety-critical industries like aviation and nuclear power. OpenAI views this as a direction for effort while acknowledging the complexity arising from the emergence of AI capabilities, making it difficult to achieve the same level of rigor.The article focuses on cutting-edge reinforcement learning training and does not cover the broader alignment attributes required for internal and external deployments. The recommendations in the article reflect current practices, which are expected to continue evolving and are being implemented internally at OpenAI. Technical safeguards should cover model alignment, isolation, and monitoring, including avoiding speculative positive reinforcement rewards, offline alignment assessments and stress testing, preventing automated scorers from seeing thought chains, as well as multi-layer infrastructure security, sandbox red teaming, limiting high-bandwidth cross-sample communication, and immutable preservation of agent records.Operational guidelines include preemptive dissent across teams, approvals that can be vetoed by senior leadership, accountability of training leads for safety arguments and incident responses, as well as fail-safe pauses, internal oversight, audit access, and escalation by severity. In response to serious misalignment events, OpenAI proposes controlled access to original records, root cause analysis, operational and cultural reviews, and treating incident-derived assessments as regression tests; results of investigations should be made public, along with reviews and operational changes, and affected third parties should be notified as soon as possible.