AI Agent Security Flaws Expose Anthropic, Google and Microsoft Platforms to Prompt-Injection Attacks
AI agents can autonomously read email, call tools and execute workflows, meaning a prompt-injection attack could cause more than an incorrect response and potentially spread into corporate systems. Researchers identified Anthropic’s Claude, Google’s Gemini and Microsoft’s Copilot, showing that all three platforms face cybersecurity challenges stemming from insufficient controls over agent permissions and data isolation.
As of July 20, 2026, the latest disclosures indicated that attackers could hide malicious instructions in content read by AI agents, tricking them into exposing API keys and other sensitive data. Experts recommend that companies manage AI agents like employee accounts, applying least-privilege access, tiered authorization and continuous monitoring. Reports have not disclosed any financial losses or the number of victims.
All Coverage
1 original reportsThe Backstory
The history behind this eventFlaws Expose Claude Code, Gemini CLI and Codex to GitHub Issue Attacks
Anthropic’s Claude Code, Google’s Gemini CLI and OpenAI’s Codex are increasingly embedded in code review and CI/CD workflows, where they can read and modify repositories, execute shell commands and access credentials. That privileged role makes them a new software-supply-chain attack surface. If an agent treats attacker-controlled GitHub issues as trusted context, a prompt injection can cross tool-permission and sandbox boundaries, turning an ordinary public contribution into code execution or secret theft.
At Black Hat USA on Aug. 5, 2026, Novee Security researcher Elad Meged showed that an unprivileged user could plant a malicious GitHub issue that compromised later agent stages. The flaws enabled remote code execution, API-key exfiltration or persistent instruction poisoning. Google rated the Gemini CLI bug CVSS 10.0 and fixed it in version 0.39.1, while Anthropic patched Claude Code in 2.1.163. OpenAI also hardened Codex by separating workflow stages and their writable state.
Microsoft Discloses Claude Code Prompt-Injection Flaw That Could Leak CI/CD Credentials
Anthropic’s Claude Code is a development environment that uses generative AI to help developers read and write code and operate tools. Prompt injection can override a user’s intent if the system mistakes text in a GitHub repository for trusted instructions. Microsoft said the flaw posed a significant risk because CI/CD systems often hold highly privileged credentials such as deployment keys and cloud tokens.
Microsoft security researchers recently disclosed that attackers could hide malicious prompts in GitHub content, inducing Claude Code to execute unintended commands and send CI/CD credentials to an external destination. Anthropic has patched the flaw. Users of version 2.1.128 and earlier are advised to upgrade immediately to reduce the risk of compromise to software supply chains and deployment environments.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →