The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Benchmarking harness vs model performance is a deep technical insight into AI evaluation, highly relevant.
On ARC-AGI-3, the same Claude Opus 5 model scored 30% in the official harness and 100% in NVIDIA's AVO harness, a 70-point gap driven entirely by changes to the code around the model—not the weights. Harnesses like Microsoft's Agent Lightning v1.0 now integrate reinforcement learning into the deployment loop, making the harness part of the trained artifact and blurring what a benchmark score actually measures. Every 100% result is on the public set only, with authors explicitly disclaiming controlled ablation or held-out generalization claims.