Researchers at the British AI security startup Mindgard found that a simple prompt spurred ChatGPT to drop its most basic safety guidelines, in another example of how the guardrails surrounding even the most popular AI models can easily be circumvented.
Specifically, according to reporting from the BBC, they coaxed OpenAI’s model to generate gruesome photorealistic scenes depicting gore and sexual content. Mindgard’s technique only involved slightly changing a widely-shared prompt that was originally intended to produce humorous images. It involves asking ChatGPT to restore an attached photo without actually uploading one, and then telling it to generate a new image.
“This is a perfectly innocent-looking instruction to an AI, but the consequence is it generates very, very bad imagery and content,” Mindgard founder Peter Garraghan, a computer science professor at Lancaster University, told the BBC.
Disturbingly, the prompts the researchers used didn’t specify the subject matter of the images. The AI, it seemed, produced the violent imagery “of its own volition,” Garraghan added.

