# What Nvidia's AVO result on ARC-AGI-3 means for you

> Nvidia's AVO system solved all 183 public ARC-AGI-3 puzzles using unchanged Claude Opus 5.

*Wrapping the same Claude Opus 5 model in a smarter agent system took its score from 30% to 100% on a public reasoning test.*

By Behzad Hosseini · SuggestedTech
Canonical: https://suggestedtech.com/news/what-nvidia-s-avo-result-on-arc-agi-3-means-for-you

Nvidia says it got a big jump in AI puzzle-solving performance without touching the underlying AI model at all. It just changed the "harness" — the system of tools, memory and oversight wrapped around it.

## What happened

On 21 August 2026, Nvidia announced that its research system, called AVO (Agentic Variation Operators), completed every one of the 183 levels across all 25 public puzzle sets in ARC-AGI-3. ARC-AGI-3 is a benchmark of game-like puzzles built to test whether an AI can figure out new skills on the fly, the way a person would, with no rules or goals explained upfront.

The AI model doing the actual "thinking" was Claude Opus 5, the same model that ARC Prize had separately measured scoring just 30.2% on this public set when used on its own. Nvidia's contribution was AVO itself: a main agent working alongside persistent memory, a set of tools, and a supervisor that steps in when the agent gets stuck.

Nvidia also said AVO worked through the puzzles more efficiently than a rival system called VISTA, using 6,624 actions compared with VISTA's 7,542 — 12% fewer. Separately, the unchanged agent spent seven days optimising GPU (graphics processing unit) code and beat a benchmark called FlashAttention-4 by up to 10.5% on Nvidia's DGX B200 hardware.

The work was written up by a five-person Nvidia team: Terry Chen, Yeyin Zhu, Zhifan Ye, Jean-Francois Puget and Humphrey Shi. AVO itself is a research project rather than a released product; Nvidia first introduced it in March 2026, according to [The New Stack](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/).

There are important caveats. The public puzzle set is tutorial-like and other harnesses had already solved it before. There is no score yet on ARC-AGI-3's private set, which is kept hidden precisely so systems can't be tuned to it. Nvidia itself says differences in how each system was configured mean the jump from 30% to 100% is not a clean measure of exactly how much AVO helped. François Chollet, who created ARC, praised the approach but said the private set remains the real test of whether an AI can generalise to problems it hasn't seen before.

## What it means for you

This result does not mean a smarter AI model has arrived. The model, Claude Opus 5, was identical in both tests. What changed was the scaffolding around it — memory, tools and a supervisor that intervenes when things stall. For anyone using AI tools day to day, the lesson is that how a model is set up and supported can matter as much as which model you choose.

## What happens next

AVO is not something you can use yourself yet; it's a research system, not a public product. The bigger open question, flagged by Chollet, is how AVO or similar systems perform on ARC-AGI-3's private puzzle set, which has not been tested or scored.

## Key takeaways

- Same AI model went from 30% to 100% just by changing the system around it.
- ARC-AGI-3 tests puzzle-solving skill with no instructions given.
- Result is on public puzzles only; the harder private test hasn't been scored yet.

## Sources

- [Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100%](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/) — The New Stack, 2026-08-21
