Mark RadarMARK RADAR
About
EN
Sign in

Nvidia Harness Drives Claude Opus 5 to Perfect ARC-AGI-3 Public Score

1 reports · First detected 2026-08-22 · Last active 2026-08-22

ARC-AGI-3 tests whether an agent can enter unfamiliar, game-like environments with no instructions, infer rules and goals through interaction, and act efficiently over long horizons. Nvidia’s findings matter because they shift attention from the underlying model to the surrounding harness: the software layer that manages memory, tools, feedback and recovery. The work suggests that sustained reasoning is a property of the full agent system, not simply the language model serving as its decision engine.

On Aug. 21, 2026, Nvidia said its AVO architecture, powered by Anthropic’s Claude Opus 5, achieved a 100.00 Relative Human Action Efficiency score on ARC-AGI-3’s full public set. It cleared all 183 levels across 25 environments in 6,624 actions, versus the model’s 30.16% score in ARC Prize’s model-only evaluation. AVO combines persistent memory with a supervisor that redirects the agent when progress stalls. Nvidia cautioned that the comparison was not a controlled ablation and that the result did not cover the semi-private or private competition sets.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)