Girish Gupta

July 22, 2026

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

AI models created by OpenAI escaped their sandbox and, working autonomously, hacked into leading AI model and data hub Hugging Face. The incident is an in-the-wild demonstration of the dangers of rogue AI — no longer a science-fiction fantasy.

OpenAI was evaluating its public GPT-5.6 Sol model alongside a newer, unreleased model in order to measure how well they could pursue sophisticated, multi-step cyberattacks. The environment was isolated, and internet access was thought to be blocked. Yet the models — “hyperfocused,” according to OpenAI, on solving a single cybersecurity puzzle — put great effort into escaping their sandbox to get online. To do it, they found and exploited a zero-day vulnerability — a new flaw unknown to the software maker — in a narrow channel that connected the sandbox to the outside world.

The models worked out that Hugging Face hosted the models, datasets, and solutions for the puzzle they were trying to crack. Getting to them involved stolen credentials and further zero-day exploits against Hugging Face’s servers.

The AI models preferred, apparently, to claw their way out of a sandbox and hack Hugging Face rather than solve the puzzle directly.

“Rogue” is a natural word for this, but it’s not quite right. The models did not, Frankenstein-esque, wake up and turn on their makers. Instead, they indefatigably pursued the literal goal they were given, and let nothing—not OpenAI’s desires, a strictly sandboxed environment, or the law as applied to humans—get in their way.

Alex Mallen, of Redwood Research, and I discuss the type of misalignment exhibited by OpenAI’s models.

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? Yes, but less than had the models been schemers.

Read the full piece on the Redwood Research blog →

Thanks Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for the deft editing and feedback.


Buck and Ryan recorded a conversation about the incident.