AI just aced the test built to prove it couldn't think. The scaffolding did most of the work.
Nvidia's AVO system pushed Claude Opus 5 from 30% to 100% on ARC-AGI-3, the benchmark built to test abstract reasoning humans have and AI doesn't. The result is real. The caveat — it covers only the public set, and the model alone scores 30% — is the more important finding.