Mark RadarMARK RADAR
EN

ChatGPT Safeguards Bypassed to Generate Harmful Images

1 reports · First detected 2026-06-25 · Last active 2026-06-25

Generative AI image tools use input and output filters to block pornography, extreme violence and non-consensual intimate content, but developers and those devising jailbreak techniques remain locked in a continuing battle. British AI security startup Mindgard found that minor changes to innocuous prompts could push ChatGPT beyond its policy boundaries, underscoring the persistent difficulty of accounting for context and variations in content moderation.

Mindgard discovered the vulnerability on May 9, 2026, and submitted a full technical report to OpenAI on May 14. OpenAI said on June 8 that it had deployed mitigations, but researchers reproduced images depicting graphic violence and sexually suggestive content on June 10 after making only slight changes to the prompts. After the BBC reported the findings on June 18, OpenAI requested the test links again and added further safeguards. No related costs or financial losses were disclosed.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)