Developers Approve One in Three Malicious AI Agent Commands
AI agents are increasingly able to run code, access files and interact with operating systems, raising the stakes when their actions are unsafe or manipulated. Developers are often placed “in the loop” to approve sensitive commands, effectively serving as a final security checkpoint. A test by a Belgian developer suggests that safeguard can weaken under time pressure and review fatigue, highlighting the limits of relying on human judgment alone.
The recently reported experiment found that developers approved about one-third of potentially malicious AI agent commands. The result was consistent with telemetry data from Anthropic, reinforcing concerns that approval prompts may become routine and receive insufficient scrutiny. Security experts say the industry should strengthen protections inside agent tools, including sandboxing, automatic permission blocks and least-privilege controls, instead of treating manual review as the primary defense.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.