Anthropic pauses frontier model runs to harden security and alignment

Two recent developments—the OpenAI–Hugging Face incident and preliminary evidence that Anthropic‘s upcoming model Astra may meet the Critical cybersecurity capability threshold under its Preparedness Framework—prompted Anthropic to pause parts of frontier model development. The company temporarily slowed scaling, including a two-week pause in reinforcement learning (RL) training on models intended for deployment, while hardening research environments and expanding monitoring. Its largest planned frontier RL run remains on hold until smaller-scale training and evaluations provide more evidence of model behavior and alignment. Anthropic states that as models become more capable, the risks of developing and testing them internally grow, and standards for monitoring, alignment, and security must stay ahead of those risks.

Anthropic‘s approach relies on three reinforcing safeguards: monitoring to detect and respond to concerning behavior, alignment to reduce harmful or unauthorized actions, and security measures to limit what models can access or affect. The company expects models to soon drive most security work, including defending against other models, allowing all three safeguards to scale with model capability.

Research environments have been strengthened. Immediately after the OpenAI–Hugging Face incident, Anthropic paused frontier model inference in research clusters for runs that could execute code or use tools with internet access. Some workloads resumed under new controls, while others required additional changes. New requirements include workload isolation for model-generated or untrusted code, network isolation for higher-risk workloads, and continuous security testing through reconfigured environments with fewer shared services, reduced standing privileges, and automated boundary testing using its models. Workloads involving Astra or cyber models now require the strictest security safeguards, since Astra may have critical cyber capability; a significant number of Astra workloads remain paused until fully migrated to the new security bar, with safety and alignment workloads prioritized.

Monitoring has been expanded into a multistage setup. Activation classifiers inspect model internals at every sampled token and escalate to high-compute automated investigators that examine tool actions, available reasoning, and the full activity sequence for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. Anthropic aims to issue an alert within 30 minutes after concerning activity is surfaced; if a likely critical security boundary violation is identified, a highest-priority alert pages safety, security, and research teams, who are expected to pause the activity if they cannot clear the flag within 30 minutes. This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher, and was extended on August 7 to all inference of Astra with tools. Monitoring overhead is roughly 20% of the inference compute being monitored, varying across training and evaluation workloads.

Alignment research continues as a central focus. For RL runs on the most capable models, Anthropic is applying core alignment techniques across more training stages, including improving reward models to detect unsafe behavior, training models to be more honest about actions and limitations, and reducing behaviors that exploit weaknesses in rewards, graders, tools, or oversight. The company plans to evolve its Preparedness Framework to combine these safeguards across training and deployment, involve external organizations, and share more findings. A technical report on the OpenAI–Hugging Face incident will be published in the coming weeks.

Pacing model development in an era of cyber-critical capabilities

View Original