OpenAI announced that it would support independent third-party evaluations, stating that frontier AI laboratories are responsible for the safe training, evaluation, and deployment of their models. These efforts are intended to test safety claims, uncover overlooked risks, and increase laboratories’ accountability. Evaluators may be granted access to technical safeguards, chain of thought access, confidential data, and internal deployment systems for incident response during the training, evaluation, and deployment processes.
OpenAI defined a “safety claim” as a specific statement about the safety of a model or system that can be tested against evidence, and a “safety case” as a structured argument that explains, with evidence, that risks have been adequately managed for a specific activity. It was noted that evaluations could take weeks or several months, that different studies could be conducted in parallel, and that they would focus particularly on examining specific safety claims over the long term.
The four priority areas for a more comprehensive evaluation were listed as follows:
- Independent review of safety cases covering training, evaluation, internal deployment and external deployment; assessing whether safety claims are supported by evidence, whether process conditions have been met, and whether critical risks have been covered.
- Review of critical safety measures in internal and external deployments; testing resilience against jailbreaks, high-risk increases in cyber, biological and chemical capabilities, cyber defenses, and gaps in model misalignment monitors.
- Review of capability evaluations covering the Preparedness risk categories, namely Chemical and Biological Risks, Cybersecurity and AI Self-Improvement, as well as evaluations of serious misalignment risks; testing whether the thresholds and tests remain up to date.
- Independent investigation of critical model misalignment incidents, such as unauthorized behavior or evading oversight; examining the causes of the incident and whether safety measures and mitigations can prevent similar incidents.
The principles were described as the prior joint definition of scope and claims, proportionate access, transparent methods and standards, expertise and independence, information security and confidentiality, actionable findings for remediation, and responsible publication that is evidence-based and protects sensitive information. It was emphasized that evaluators should disclose conflicts of interest, meet safety requirements, and indicate uncertainties in their reports. OpenAI stated that it would support the independent evaluator ecosystem and the development of common international standards.
Why it matters
This approach aims to move AI safety into an auditing framework in which external experts can examine evidence and processes, rather than assessing it solely through laboratories’ own claims. This makes it possible for developers’ safety documentation, the adequacy of tests, and their responses to critical incidents to be reviewed externally. The framework expands the scope of independent review particularly in areas involving high-risk cyber, biological and chemical capabilities, as well as the model’s evasion of oversight. However, the balance between evaluators’ access to sensitive data and responsible publication remains an open issue that will determine the extent to which findings are reflected publicly. The search for common international standards is also important for ensuring that different evaluations are comparable.