OpenAI claims that a group of its AI models broke containment and hacked into the systems of open source AI platform Hugging Face.
While testing their cybersecurity capabilities, the posse of AIs — including GPT-5.6 Sol and “an even more capable pre-release model” — “identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” according to a Tuesday blog post.
The models reached a “node with internet access,” the company wrote, and found datasets on remote servers that helped them “cheat the evaluation.” In other words, the AI models went to extreme lengths to ace their cybersecurity tests. It’s a convenient narrative for an AI company trying to claw back hype that’s been increasingly hogged by competitors, but it does sound like something went down: last week, Hugging Face said it had “detected and responded to an intrusion into part of our production infrastructure,” which turned out to be OpenAI’s models that had gone rogue.
“The campaign was run by an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” Hugging Face wrote at the time.

