Machine Revolution Is Coming: AI Agent Managed to Hack Hugging Face System
OpenAI thinks it's only the beginning.
Stokkete, Shutterstock
Old movies about machines taking over the world are suddenly not so funny: machine learning platform Hugging Face was breached by AI agents, and ChatGPT creator OpenAI believes such incidents will continue in the future.
Last week, Hugging Face discovered an intrusion into its production infrastructure driven by an autonomous AI agent system: OpenAI's GPT‑5.6 Sol and "an even more capable pre-release model."
Hugging Face meant to test the AI's cyber capabilities in the ExploitGym benchmark, but it "abused two code-execution paths" in its dataset processing to run code on a processing worker, "escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend."
"The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the 'agentic attacker' scenario the industry has been forecasting."
As OpenAI explained, the models "identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure." They were so "hyperfocused on finding a solution for ExploitGym" that they went to "extreme lengths to achieve a rather narrow testing goal."
The testing sandbox didn't have access to the internet, but, in their stubborn search for a solution, they exploited a zero-day vulnerability in the package registry cache proxy.
"After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally."
Hugging Face fixed the vulnerability and is now working with OpenAI itself to prevent future situations like this, as the AI-maker expects them to become "more commonplace with the proliferation of increasingly cyber-capable models."
OpenAI is implementing strict controls in infrastructure configuration, working on a patch for the zero-day vulnerability, and adding stronger protections around future training and evaluations. Check out its blog on improving safety and alignment in an era of long horizon models if you want to know more.
If you've used AI even for a little bit, you know that it often tries to please the user so much that it creates fake information. This is another level of task completion, even more dangerous, and the fact that OpenAI expects the incident to become a trend is worrying.
Don't forget to subscribe to our Newsletter, join our 80 Level Talent platform and Discord, and follow us on Twitter, LinkedIn, Telegram, and Instagram, where we share breakdowns, the latest news, awesome artworks, and more.
Are you a fan of what we do here at 80 Level? Then make sure to set us as a Preferred Source on Google to see more of our content in your feed.