OpenAI’s case for global AI safety standards

OpenAI‘s post “Building standards for the next phase of AI” opens with its mission that AGI should benefit all of humanity. Following Sam Altman and Jakub Pachocki, it prioritizes three goals: navigate the next period of AI progress by building an automated AI researcher and iterating with it on alignment, deliver the scientific and economic benefits of intelligent machines, and empower everyone with a personal AGI. The post argues that safe navigation requires alignment research to keep pace with capabilities, and that shared standards to guide development across labs and countries are as important as alignment research itself.

A central concern is recursive self-improvement (RSI). Automated AI research can involve varying degrees of human supervision, and as systems take on more of the work of developing successive generations of AI, they could drive RSI even while people remain involved, potentially accelerating progress rapidly. The post states that fully autonomous RSI is not happening today and should not be pursued unless it can be done safely; proceeding without care could result in humans losing practical control over AI development and being unable to oversee research processes they no longer understand. It cites the disclosed Hugging Face Incident as a preview of risks that could become much more severe without robust safeguards and alignment.

The post then argues that international standards for safety and security practices in frontier AI development may be as important as alignment research. Standards can create shared definitions of high-quality evidence and agreed-upon baselines for the rigor of technical safeguards, helping answer the question of what good looks like in mitigating catastrophic AI risk. OpenAI points to existing institutions like the Center for AI Standards and Innovation (CAISI), state laws, and a federal framework for AI, but contends that international efforts are needed to solve three challenges: fragmentation across nations’ evaluations and reporting requirements, collective action problems where independent national decisions produce unwanted outcomes, and uneven distribution of frontier development expertise. These challenges apply to both open and closed models. The post is explicit that labs pursuing automated research must take accountability for safety, and that standards are rooted in avoiding concentration of power and producing better outcomes.

On implementation, OpenAI proposes a mechanism that facilitates complementary national and international frontier standards. One approach is to leverage the emerging network of AI safety institutes, existing in Australia, Canada, Germany, France, Kenya, Japan, Korea, Singapore, India, and the United Kingdom, to facilitate standard setting through CAISI and national industry bodies, focusing on frontier AI models and developers as measured by capability benchmarks, and on benefit-risk management for automated AI research including RSI. The United States can build on CAISI‘s creation of the International Network for Advanced AI Measurement, Evaluation, and Science in 2024. Crucially, the standards would not be licenses, mandatory prerelease review, or approval requirements; national governments would decide how to incorporate them into their legal systems. The process should consult open and closed model developers, independent experts, and academia, and be designed so it does not advantage particular companies, countries, or business models, including making it harder for new entrants or open-weight developers to compete. Lessons can be drawn from aviation and financial stability, and collaboration should include bodies such as ISO, the Frontier Model Forum, Agentic AI Foundation, the Open Secure AI Alliance, and implementation-focused organizations like the Appia Foundation.

For RSI specifically, OpenAI calls for common measurements and incident reporting protocols, including evaluation of RSI-relevant progress and the amount of autonomous research happening within an AI company, human oversight triggers for automated research, and incident classification, tracking, reporting, and response frameworks. Its research acceleration report and misalignment reporting framework are cited as early contributions. The post also urges critical infrastructure operators and governments worldwide to establish secure communication channels for sharing national security concerns and vulnerabilities, and describes dialogue between the United States and China as a positive step.

Finally, the post argues the United States should lead the effort to develop global technical standards. Pacing is not about maintaining a predetermined speed; technically, it is about ensuring alignment research and deployment stay ahead of capabilities, and societally, it is about using standards and institutions to produce positive-sum outcomes. The United States is well positioned because its AI industry is at the technical frontier and it holds a privileged global network position in finance, trade, defense, technology, and information systems. Geopolitical competition will involve AI adoption and diffusion, and leading now may determine whether the US shapes the global AI framework or watches a fragmented, uneven, and conflict-ridden system take hold. Strong national governance connected through practical international cooperation is presented as a path to stronger safeguards, continued innovation, and broad access to AI benefits.

Building standards for the next phase of AI

View Original