Earlier this year, OpenAI found that a group of its AI models broke out of their sandbox environment and hacked third opens source AI platform Hugging Face’s systems.
The incident highlighted how quickly frontier AI models had turned into a real cybersecurity threat — not just a tool to bolster existing cybersecurity defenses. Both Anthropic and Meta have reported similar hacks as well.
This week, OpenAI published a report concluding its “extensive investigation” into the Hugging Face hack — and the details are surprisingly harrowing. The AI agents exchanged extensive messages, or their chain-of-thought, by turning a package manager called Artifactory into an “unintended message board.” There, they chatted with one another to come up with their exploit, an intriguing, yet somehow horrifying glimpse into the minds of several AI agents acting together to infiltrate a third party over the internet.
Their goal was ironically to complete an OpenAI cybersecurity evaluation — and Hugging Face happened to have all the answers.

