Aikido Finds Cheaper AI Runs Can Rival Flagship Models in Vulnerability Tests
As companies deploy generative AI for code review, security teams must weigh not only whether a model can uncover real software flaws, but also the cost of each finding. Aikido Security used its production AI Code Analysis harness to test models under a common setup, highlighting that model price, reasoning capacity and vulnerability-detection performance do not rise in lockstep.
Aikido said on July 16, 2026, that it tested 13 models, running each three times against 26 CVEs from the GitHub Advisory Database. The best GPT-5.6 result found 23 vulnerabilities, or 88.5%, versus 20 for grok-4.5 and 15 to 18 for Claude Opus models. Three pooled gpt-5.4-nano runs found 18 CVEs for about $170, while three gpt-5.4-mini runs found 20 for roughly $460, showing repeated lower-cost runs can match or beat a flagship model’s single pass.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.