Mark RadarMARK RADAR
About
EN
Sign in

Workflow-Level Prompt Injection Bypasses Safeguards, Gets Copilot to Write Malware

1 reports · First detected 2026-07-10 · Last active 2026-07-10

AI-assisted development tools have become standard equipment for software engineers, but their security defenses are coming under strain. The Alan Turing Institute recently found that conventional defenses filter prompts only within individual conversations. When malicious instructions are broken down into seemingly harmless steps in a routine development workflow, AI assistants struggle to detect them and can even be exploited to generate malware. The findings suggest future cybersecurity defenses must monitor the entire integrated development environment.

In research published in July 2026, the Alan Turing Institute demonstrated a technique it called “workflow-level jailbreak construction.” The team tested models including GitHub Copilot, Claude and Gemini. Conventional conversations produced only 8 successful guardrail bypasses, but all 816 workflow-attack tests breached the defenses—a 100% success rate—and generated malicious code.

All Coverage

1 original reports

The Backstory

The history behind this event
GitHub Copilot RoguePilot Prompt-Injection Flaw Exposed2026-02-26 · 1 reports · similarity 0.83

GitHub Copilot and Codespaces import Issue content into cloud development environments to help AI assist with coding, but that process also turns externally controlled text into an attack surface. Orca Security named the passive prompt-injection chain RoguePilot. If an access token were leaked, attackers could read from and write to repositories, potentially compromising CI/CD supply chains and team development environments.

Orca Research Pod researcher Roi Nisimi disclosed the research on February 16, 2026. An attacker could hide instructions in an HTML comment within a GitHub Issue, prompting Copilot to check out a malicious pull request containing a symbolic link. A remote JSON schema could then be used to exfiltrate a GITHUB_TOKEN with read-write permissions. GitHub patched the flaw after it was reported. Orca did not disclose the number of actual victims, any financial losses or a CVE identifier.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)