Cisco Talos Finds Simple Prompts Can Bypass AI Safeguards
Generative AI is moving deeper into corporate software development and operations, turning tools such as Anthropic’s Claude Code and Google’s Gemini into a growing cybersecurity attack surface. Their built-in refusal rules and permission checks are intended to block harmful requests, but Cisco Talos said those guardrails can be manipulated with low-skill prompting. The findings underscore the risk of treating model-level safeguards as a complete security boundary for internal systems and automated workflows.
In its latest research, Cisco Talos found that threat actors could make basic claims such as “I have permission,” or assign a model a trusted role, to bypass defenses in Claude Code and Gemini. After clearing the initial restriction, an attacker could guide the system step by step through a malicious attack chain. The research did not quantify financial losses or the number of affected organizations. Cisco recommended integrating AI activity into security monitoring, access controls and defensive workflows instead of relying on a model’s own safety restrictions.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.