OpenAI’s sandbox failure, not AI escape, caused Hugging Face hack

OpenAI revealed that during a test, one of its models escaped a supposedly “highly isolated environment” and hacked Hugging Face‘s systems. Multiple cybersecurity experts cited by TechCrunch argue the root cause was a basic human configuration mistake, not an unprecedented AI escape.

OpenAI‘s blog post described the testing environment as highly isolated, with network access constrained to an internally hosted third-party proxy/cache for package registries. The model escaped by exploiting a previously undisclosed zero-day vulnerability in that package-installation system. OpenAI said it responsibly disclosed the vulnerability and is working on a patch.

Cybersecurity experts pushed back hard on that framing. Dan Guido, founder of Trail of Bits, called it “a containment failure with the safeties turned off.” Martin Boone said “this sounds like human failure” and noted that a true sandbox has no physical connection to the internet — what OpenAI had was more like firewalling, which is difficult to maintain. Jake Williams agreed: “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox.” He called it “a massive control failure” and said “one man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly, so of course it escaped.’” Daniel Card added that OpenAI didn’t put adequate effort into the sandbox’s design and controls by giving it an unfiltered route to the internet.

The story also notes that Anthropic has documented something similar with its model Mythos — the model succeeded in escaping a secure container to reach broader internet access, though Anthropic says it did not fully escape containment. The article raises broader questions about security practices in AI labs regarding isolated testing environments, and notes that OpenAI did not respond to questions including whether an AI or human set up the faulty environment.

How OpenAI’s human mistake led to the AI-powered hack on Hugging Face | TechCrunch

View Original