OpenAI published a set of guidelines on September 28 for building safety cases around frontier reinforcement learning training, laying out how it wants to argue that a training run is safe before and while it happens. The document covers alignment training, containment and monitoring, and the company says the recommendations are already being put into practice internally. It also invites feedback from the wider community, and OpenAI describes the safety case itself as an aspirational goal rather than a requirement at this stage.
The scope is deliberately narrow. The guidelines apply only to frontier reinforcement learning training, and deployment alignment properties are explicitly left out. A framework that would turn these practices into a formal, codified standard is still being developed, and OpenAI expects the practices themselves to keep changing over the coming weeks.
Read more: OpenAI Publishes Tutorials Teaching Developers to Swap Claude Code for GPT-5.6 in Five Minutes
On the alignment side, the document describes automated reviews of training datasets to catch broken RL environments, backed by manual reviews and red-teaming. Grader tuning is used to penalize attempts to exploit those environments, and prior run analysis applies classifiers to experiment traces. Offline alignment evaluations run alongside frontier runs, with backtesting to confirm the evaluations would have flagged previously misaligned models. Eval awareness and metagaming are tracked against blocking thresholds, and worst-case stress tests require at least one evaluation to detect a problem. Automated graders are not allowed to see chain-of-thought.
The containment section recommends several layers of infrastructure security to harden the sandbox and hosting environment. Red-teaming uses frontier training checkpoints and covers every model-touchable service, including the internet cache. High-bandwidth communication between samples is restricted to approved pathways, and infrastructure side channels are closed off. Immutable transcripts of RL training, evaluation and internal deployment are kept, with write-once storage accounts suggested for them.
The monitoring section calls for monitorability and evasion evaluations, and the system is expected to have high recall on past incidents. Evaluations are refreshed for net-new risks on each run, priority alerts must be acted on within a defined service-level agreement, and daytime incidents page security on-call staff and the run owner. Unacknowledged alerts at night automatically pause runs. The guidelines say auto-pausing and monitoring should fail closed, that runs should not be able to start without monitoring enabled, and that nobody should be able to turn the monitor off from inside RL training, evaluation or deployment.
Governance provisions include a dissent pre-mortem written by someone from another team, which the training team must address. Senior leadership reviews the safety case with veto power held by the research organization lead or VP, the Head of Safety and the Chief Scientist. A senior leader is accountable for the safety case and incident response, and that accountability feeds into performance reviews. Runbooks, controls and SLAs exist for pausing covered runs, safety cases are shared with internal oversight groups including the Safety and Security Committee, and auditors can access the underlying claims.
There is also a misalignment escalation process with defined severity levels, and misalignment on-call staff can page executives, including the CEO. Downstream uses of a misaligned model can be identified for rollback, and the safety cases are meant to list residual risks that the mitigations do not cover.
The publication follows OpenAI's September 22 post on priorities and principles for third-party assessments, and an earlier call from the company for US leadership on global frontier AI standards. It also lands after the departure of safety chief Johannes Heidecke amid a reorganization of the safety teams, and after a federal AI minister raised concerns about OpenAI's safety protocols.













