Nvidia's AVO agent aces ARC-AGI-3 — all 183 levels
Nvidia published research on its Agentic Variation Operators (AVO) system, which wraps a frontier reasoning model with persistent memory and dynamic tool use to achieve a perfect score on ARC-AGI-3, the hardest public general-reasoning benchmark. The same system also generated GPU kernels that outperform FlashAttention-4 on DGX B200 hardware without any human input.
- AVO elevated Claude Opus 5 from a 30% baseline to 100% across all 183 ARC-AGI-3 levels
- In a parallel engineering task it produced kernels with up to 10.5% higher throughput than FlashAttention-4
- The result is a harness architecture, not a new model — it wraps existing frontier models with memory and tools