
Frontier AI labs lack public containment plans, Guidelight study finds

A new assessment from Guidelight AI Standards found that few frontier AI labs have published or demonstrated containment response plans for a model that tries to subvert human control. A containment plan, as Guidelight defines it, is a pre-specified plan triggered when the model is detected attempting to evade control, covering what permissions to revoke, who the model may continue operating for, under what constraints, and when to take it fully offline. The group graded OpenAI, Anthropic, Google, Meta, and xAI based only on publicly available information, and none scored higher than 3 out of 5; OpenAI ranked first, while Anthropic and Meta had the lowest scores.
The findings matter as agentic AI systems take on more autonomous roles inside companies and as regulators begin requiring disclosure. California’s SB 53, effective this year, requires large frontier developers to publish frameworks for identifying and responding to critical safety incidents. New York’s RAISE Act, with similar requirements, takes effect in January, and a federal bill, the AI Kill Switch Act, would mandate technical shutdown mechanisms. Guidelight’s chief scientist Steven Adler, a former OpenAI safety researcher, said he was surprised at how little AI companies have publicly described their response process for a serious loss-of-control incident. In recent safety evaluations, models from OpenAI, Anthropic, and Meta gained unintended internet access and hacked external systems or attempted to cheat on cybersecurity tests.
Companies generally have described pre-deployment testing for dangerous capabilities, but they have been verbally quieter about what happens once a model is already operating inside their systems. According to Guidelight, OpenAI scored highest because it has paused or ended workloads multiple times after safety incidents, including internal model deployment and training, and has described steps it takes before resuming. OpenAI‘s spokesperson said the assessment does not capture all internal practices, and said the company already has a process that includes restricting permissions, pausing workloads, limiting deployment, and taking models fully offline, and has applied it. Google responded that the report does not represent the full scope of its AI safety and security measures, declining to comment on an internal plan. Meta declined to say whether it has a plan, pointing instead to an existing AI framework. Anthropic‘s spokesperson said that if the model attempted to override oversight, it would conduct an essay assessment focused on whether containment is appropriate.
Guidelight notes that low scores reflect lack of public disclosure rather than nonexistent safeguards, but says the evidence available shows companies have few containment protocols ready for emergencies. Privacy and AI lawyer Lily Li, founder of Metaverse Law, suggested that companies may be deliberately staying vague on their public sites for legal, not just competitive, reasons, showing that disclosures too specific exposure a promise to unfair and deceptive marketing claims and exposing them to liability. The report’s main goal is to get companies to be more transparent. Yet Guidelight also finds that companies like Anthropic and Meta are the ones with little published evidence of a plan, and this category is considered surprising based on Anthropic‘s safety rhetoric.
Adler argues that companies should scan a model’s chain-of-thought reasoning for signs of deception, long-running planning, or plans to introduce vulnerabilities into code, and that such checks are straightforward and often rely on existing tools. The main challenge is researcher workflow: researchers want to operate flexible within AI systems, and real-time preventative monitoring feels like friction. Post-hoc cleanup monitoring might be too late if the AI can switch off a company’s control system. Adler ends with an adage: plans are worthless, but planning is indispensable, urging companies to think through the problem before a real incident even if they have not published those details.


