OpenAI sets priorities and principles for third-party AI safety assessments

OpenAI published a statement on how it intends to work with independent third-party assessors on frontier AI safety, proposing four priority areas and eight principles for rigorous, secure, and independent assessments. The post defines two recurring terms: safety claims, meaning the testable assertions a lab makes about the safety of its training, evaluation, and deployment; and safety cases, the broader body of evidence that substantiates those claims.

The first priority is independent assessment of safety cases across training, internal deployment, and external deployment. OpenAI says these assessments require expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse, and red teaming. Multiple assessors may cover different parts of a case. Key questions include whether evidence for safety cases is substantiated, whether the conditions of each safety case were actually followed, whether the most urgent risks are covered, and whether training incentives that could reward deception, reward hacking, destructive actions, or circumvention of restrictions are identified and reduced.

The second priority is assessment of critical safeguards across internal and external deployments. These safeguards include model-level protections, enforcement safeguards, security safeguards, and misalignment monitors. OpenAI says technical partnerships can expose weaknesses now, improve assessment methods, and accelerate standards development, especially for internal deployments where safety and security standards are still nascent. Example questions cover robustness to jailbreaks under grey-box access, capability uplift in high-risk domains like cyber and bio, how agents interact with cyber defenses under realistic conditions, gaps in misalignment monitors that could lead to loss of control, reliability of chain-of-thought monitoring as evidence, whether monitoring is hard to disable, and whether safeguards match model capabilities.

The third priority is assessment of capability evaluations covering Preparedness risk categories — chemical and biological risks, cybersecurity, and AI self-improvement — plus alignment evaluations for severe misalignment risks. OpenAI notes that as evaluations saturate, they need to be refreshed to keep measuring more advanced capabilities. Relevant questions include whether evaluations match the Preparedness risk threshold definitions, whether thresholds are set correctly, whether new tests meaningfully measure advanced capabilities, and what behaviors alignment evaluations may miss.

The fourth priority is independent investigation of critical misalignment incidents, including models acting without authorization or evading oversight. OpenAI references the Hugging Face incident as an example where an independent third party can be useful. Investigators need cyber forensics expertise, alignment expertise, large-scale chain-of-thought analysis, and sufficient staff and resources. Some details may be sensitive because incident response can involve internal and third-party data. Findings should identify what behavior occurred, the contributing factors, and whether safeguards and remediation would prevent similar incidents.

The post then lists eight principles for effective assessments. First, scope and safety claims should be mutually agreed and pre-registered, and reports should state what was and was not assessed. Second, assessors should receive proportionate access within legal, security, and IP constraints, potentially through designated representatives or privacy-preserving mechanisms. Third, methodology should be transparent, using established standards where available, and reports should distinguish direct findings from interpretation. Fourth, assessors should have relevant expertise and disclose conflicts of interest, with safeguards such as recusal or exclusion periods. Fifth, security and confidentiality protections should match the sensitivity of accessed systems and information, and access on company-managed devices or premises may be appropriate. Sixth, findings should be actionable, with a reasonable remediation period before publication where appropriate, and lessons should be shared for developers, deployers, and defenders. Seventh, publication should be as open as possible while protecting sensitive information, with confidential reporting to governance bodies when needed and principled redaction policies. Eighth, assessors should maintain editorial independence while ensuring confidentiality and IP protections.

OpenAI says the independent evaluation ecosystem is still growing and that no single third party can comprehensively cover all frontier safety questions. It commits to helping the ecosystem grow through deeper access across training, evaluation, and deployment, and to supporting shared international standards through future laws and private governance institutions. The document also notes that these priorities and principles focus on private and non-profit assessment organizations, complementing separate government testing arrangements with distinct roles and responsibilities.

Priorities and principles for effective third party assessments

View Original