Nvidia
What Nvidia's AVO result on ARC-AGI-3 means for you
Wrapping the same Claude Opus 5 model in a smarter agent system took its score from 30% to 100% on a public reasoning test.
The answer
Nvidia's AVO system solved all 183 public ARC-AGI-3 puzzles using unchanged Claude Opus 5.
Nvidia says it got a big jump in AI puzzle-solving performance without touching the underlying AI model at all. It just changed the "harness" — the system of tools, memory and oversight wrapped around it.
What happened
On 21 August 2026, Nvidia announced that its research system, called AVO (Agentic Variation Operators), completed every one of the 183 levels across all 25 public puzzle sets in ARC-AGI-3. ARC-AGI-3 is a benchmark of game-like puzzles built to test whether an AI can figure out new skills on the fly, the way a person would, with no rules or goals explained upfront.
The AI model doing the actual "thinking" was Claude Opus 5, the same model that ARC Prize had separately measured scoring just 30.2% on this public set when used on its own. Nvidia's contribution was AVO itself: a main agent working alongside persistent memory, a set of tools, and a supervisor that steps in when the agent gets stuck.
Nvidia also said AVO worked through the puzzles more efficiently than a rival system called VISTA, using 6,624 actions compared with VISTA's 7,542 — 12% fewer. Separately, the unchanged agent spent seven days optimising GPU (graphics processing unit) code and beat a benchmark called FlashAttention-4 by up to 10.5% on Nvidia's DGX B200 hardware.
The work was written up by a five-person Nvidia team: Terry Chen, Yeyin Zhu, Zhifan Ye, Jean-Francois Puget and Humphrey Shi. AVO itself is a research project rather than a released product; Nvidia first introduced it in March 2026, according to The New Stack.
There are important caveats. The public puzzle set is tutorial-like and other harnesses had already solved it before. There is no score yet on ARC-AGI-3's private set, which is kept hidden precisely so systems can't be tuned to it. Nvidia itself says differences in how each system was configured mean the jump from 30% to 100% is not a clean measure of exactly how much AVO helped. François Chollet, who created ARC, praised the approach but said the private set remains the real test of whether an AI can generalise to problems it hasn't seen before.
What it means for you
This result does not mean a smarter AI model has arrived. The model, Claude Opus 5, was identical in both tests. What changed was the scaffolding around it — memory, tools and a supervisor that intervenes when things stall. For anyone using AI tools day to day, the lesson is that how a model is set up and supported can matter as much as which model you choose.
What happens next
AVO is not something you can use yourself yet; it's a research system, not a public product. The bigger open question, flagged by Chollet, is how AVO or similar systems perform on ARC-AGI-3's private puzzle set, which has not been tested or scored.
Sources
- Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100% — The New Stack, 21 August 2026