Anthropic Redeploys Claude Fable 5 After Lifting of Export Controls

On June 12, the US government imposed export controls on Anthropic‘s Claude Fable 5 and Mythos 5, requiring immediate access restrictions. Unable to verify nationality in real time, Anthropic suspended both models for all users. The controls were lifted on June 30. Fable 5 becomes available July 1 globally on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork (pro, max, team, and select enterprise plans get up to 50% of weekly usage limits through July 7, then usage credits). AWS, Google Cloud, and Microsoft Foundry access will be re-enabled as soon as possible. Mythos 5 access has been restored for a set of US organizations following government approval on June 26.

The export control directive followed a report from Amazon researchers describing a technique to bypass Fable 5’s safeguards: prompting it to identify software vulnerabilities, and in one case produce exploit code. Anthropic‘s testing found that many less capable models (Claude Opus 4.8, GPT-5.5, Kimi K2.7, and others) could identify the same vulnerabilities and produce the same exploit demonstration. The behavior was a borderline case for Fable 5’s safeguards, involving routine defensive cybersecurity work, not unique Mythos-level capabilities. Anthropic worked with the government to train an improved safety classifier that blocks the reported technique in over 99% of cases. The classifier may also block benign requests more often during routine coding, and Anthropic plans to refine it.

Anthropic describes its defense-in-depth approach: multiple safety mechanisms including classifiers that set a large safety margin—blocking many benign requests to reduce the chance of missing harmful ones. The safety margin also mitigates narrow jailbreaks, which only unblock minor behaviors. No universal jailbreaks for Fable 5 have been discovered. The company acknowledges that perfect robustness is impossible and expects continued minor jailbreaks.

To address the lack of industry consensus on jailbreak severity, Anthropic is partnering with Amazon, Microsoft, Google, and other Glasswing partners to draft a framework. The proposed severity scoring uses four criteria: capability gain (how far beyond existing tools), breadth of capability gain (number of distinct offensive tasks), ease of weaponization (human effort to turn into attack), and discoverability (how easy to obtain the technique). Responses would be calibrated to severity, with the most severe cases triggering immediate mitigations and 24/7 monitoring.

Anthropic also details deeper collaboration with the US government on frontier AI security: pre-release government access and evaluation for models advancing national-security-relevant capabilities, rapid information sharing on safeguards, dedicated resources for joint research, and working toward a common industry security standard. The company calls for these rules to be codified in strong regulation applied equally across frontier model developers.

Redeploying Claude Fable 5

View Original