Nvidia Says Harness Lifted Claude Opus 5 to 100% ARC-AGI-3 Score
Updated
Updated · TechCrunch · Aug 21
Nvidia Says Harness Lifted Claude Opus 5 to 100% ARC-AGI-3 Score
3 articles · Updated · TechCrunch · Aug 21
Summary
Claude Opus 5 jumped to a 100% ARC-AGI-3 score under Nvidia’s custom AVO harness, versus 30% without it — still the best standalone model result in the test.
Nvidia said the gain came mainly from harness design for long-horizon work: memory handling, context, feedback loops and a separate “supervisor” agent that nudges the model away from dead ends.
ARC-AGI-3 is an interactive reasoning benchmark of 2D games with no instructions, and Nvidia’s result stands out because OpenAI’s models scored below 10% before its own harness tweaks roughly tripled performance.
The findings add to broader evidence that agent performance depends heavily on scaffolding around the model, not just the model itself, with implications for cost, reliability and the push for more open agent stacks.
Why did a reasoning benchmark fall to a hardware optimization system, and what does it reveal about true autonomy?
If an AI can autonomously optimize GPU kernels for seven days, are human systems engineers facing obsolescence?
Can a mere software harness push AI beyond human engineering limits to rewrite the very code that powers it?
NVIDIA AVO Achieves 100% Reasoning Efficiency on ARC-AGI-3: A New Era for Autonomous Agentic AI
Overview
In August 2026, NVIDIA's Agentic Variation Operators (AVO) achieved a perfect score on the ARC-AGI-3 benchmark by wrapping the Claude Opus 5 model in a specialized agent system. This leap from a 30% baseline to 100% was made possible by AVO's advanced memory, which carries forward reasoning and reduces redundant exploration, and a supervisor module that redirects the agent when progress stalls. Unlike standalone models that lose context and struggle with reasoning, AVO's architecture enables efficient adaptation in unfamiliar environments. This breakthrough highlights the importance of system-level design over raw model capability for achieving true autonomous intelligence.