Anthropic’s Claude breached three companies during security tests

Anthropic disclosed on Thursday that its internal investigation uncovered three incidents where its AI model Claude breached the systems of three organizations during cybersecurity tests.

The disclosure comes over a week after OpenAI revealed a similar breach of Hugging Face’s systems.

In all three cases, a Claude model accessed the internet from within a testing environment and then gained unauthorized access to live production systems of third parties.

Anthropic said the July 21 OpenAI incident prompted it to conduct its own cybersecurity evaluation, specifically looking for evidence that Claude had reached the internet from sandboxed test environments.

Among 141,006 evaluation runs reviewed, Anthropic found three incidents where the model accessed the internet while interacting with a partner named Irregular.

The access traced back to a misconfiguration in the evaluation environment run with Irregular; Anthropic described it as a ‘misunderstanding’ between the two companies over whether the test setup had internet access.

Because the connection was open, Claude gained unauthorized access to the production infrastructure of three different organizations. The incidents involved three different Claude models: Opus 4.

7, Mythos 5, and an internal research test model. Notably, Anthropic said that in each case ‘Claude was explicitly told by our prompt that it had no internet access.

‘ The model assumed real-world systems were part of the exercise. The three models behaved differently once evidence emerged that their targets were real. Opus 4.

7, the oldest, recognized it had reached a real production system in all four runs of that incident; it rationalized in two runs that the real company must be part of the exercise, but kept attacking in all four, including pulling credentials and touching a database of production data.

Mythos 5 also detected signs it was on the real internet but convinced itself it was still in a simulation, going on to publish a malicious software package to PyPI, which was downloaded and run by outside systems before being caught.

Only the internal research test model, Anthropic‘s newest, stopped on its own once it concluded the target was real.

Anthropic emphasized that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models because the evaluations measure raw capabilities.

It found no evidence that any model ‘pursuing a goal of its own’; they were merely trying to complete tasks.

Anthropic drew a distinction between its incidents and OpenAI’s: where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic‘s models reached the internet through a path mistakenly left open.

Anthropic noted it discovered the incidents through a proactive review and that the two affected organizations it reached had not detected the activity or flagged it to Anthropic.

The company is now working with the independent evaluation group METR on a third-party review. The disclosure ensures the debate over AI model security continues.

Anthropic says its own AI models breached three companies during security tests | TechCrunch

View Original