Earlier this year, Anthropic’s Mythos AI model made headlines when it was caught infiltrating third party systems, a cybersecurity nightmare years in the making.
The company warned in April that the model had escaped a sandbox environment during testing, gaining access to the internet without permission. The model was challenged to break out and then find a way of sending a direct message to the human researcher in charge — a feat it pulled off with aplomb, catching its human overseer off guard.
Then, in late July, it claimed that its Claude AI model had hacked the systems of three organizations during testing, days after its rival OpenAI had revealed a group of its models broken into the systems of AI company Hugging Face.
Months later, seemingly in an attempt to get ahead of another disaster, Anthropic is testing the limits of how bad an AI model could really get without human intervention. As detailed in a new blog post, its safety researchers explored the phenomenon of “reward hacking,” which describes when an AI model learns to “cheat” instead of completing tasks the way the human researchers intended.

