Nvidia Harness Drives Claude Opus 5 to Perfect ARC-AGI-3 Public Score
ARC-AGI-3 tests whether an agent can enter unfamiliar, game-like environments with no instructions, infer rules and goals through interaction, and act efficiently over long horizons. Nvidia’s findings matter because they shift attention from the underlying model to the surrounding harness: the software layer that manages memory, tools, feedback and recovery. The work suggests that sustained reasoning is a property of the full agent system, not simply the language model serving as its decision engine.
On Aug. 21, 2026, Nvidia said its AVO architecture, powered by Anthropic’s Claude Opus 5, achieved a 100.00 Relative Human Action Efficiency score on ARC-AGI-3’s full public set. It cleared all 183 levels across 25 environments in 6,624 actions, versus the model’s 30.16% score in ARC Prize’s model-only evaluation. AVO combines persistent memory with a supervisor that redirects the agent when progress stalls. Nvidia cautioned that the comparison was not a controlled ablation and that the result did not cover the semi-private or private competition sets.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →