Mark RadarMARK RADAR
About
EN
Sign in

Microsoft Uncovers Poisoned AI Models Triggered by Keywords, Releases Detection Tool

1 reports · First detected 2026-04-07 · Last active 2026-04-07

Microsoft researchers found that some poisoned AI models may contain hidden backdoors. The models behave normally during routine testing and ordinary queries but generate incorrect responses or act maliciously when they receive specific trigger words. Traditional performance benchmarks may fail to detect such risks, potentially threatening the model supply chain and the security of enterprise applications.

Microsoft recently published its findings and simultaneously released a detection tool to help developers identify trigger-based anomalies in models. It also advised users to immediately stop using and inspect any AI system that responds irrationally or abnormally after encountering particular words. Available reports did not disclose the tool's release date, the number of affected models or the amount of any losses.

All Coverage

1 original reports

The Backstory

The history behind this event
After this
Microsoft’s New Cyber Model Lifts MDASH Score, Halves Costfirst seen 2026-07-28 · 1 reports · similarity 0.72 · same topic: Microsoft

Microsoft’s MDASH, short for Multi-Model Agentic Scanning Harness, is an AI-driven security system rather than a standalone model. It coordinates more than 100 specialized agents to inspect code, challenge suspected findings, remove duplicates and reproduce exploitable flaws. The approach matters because software audits are expensive and periodic, while code changes continuously. A system that can validate vulnerabilities with fewer false alarms and lower inference costs could move enterprise security toward continuous scanning and shorten the window in which attackers can exploit overlooked bugs.

On July 27, 2026, Microsoft unveiled MAI-Cyber-1-Flash, its first dedicated cybersecurity model, and said the model-powered MDASH configuration scored 95.95% on CyberGym, which covers 1,507 known vulnerabilities across 188 open-source projects. Microsoft said the model handles as much as 90% of tasks, with GPT-5.4 taking the hardest 10%. The result topped GPT-5.5 Cyber’s 85.6%, Anthropic’s Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6%, while costing 50% less than Microsoft’s previous best MDASH setup. The company did not disclose a dollar cost.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)