Skip to content

Benchmarking Opus 5 on SlopCodeBench

7.6 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
7
community
7
strategic
6
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Benchmarking Opus 5 on SlopCodeBench is directly relevant to AI coding agents and evaluation.

AI/ML github.com
Contribute to humanlayer/advanced-context-engineering-for-coding-agents development by creating an account on GitHub.
Summary

Opus 5 achieved only a 24% strict pass rate on a subset of SlopCodeBench, a new long-horizon benchmark from UW Madison that tests models on evolving codebases with undisclosed requirements across multiple checkpoints. All tested models (Opus 4.8, Sonnet 5, Opus 5) failed to complete any challenge defect-free, with Opus 5 writing five times more functions than Opus 4.8 and showing increased code smell over time. The benchmark remains unsaturated—GPT-5.4 and Opus 4.6 scored 11% and 17% respectively in the original paper—suggesting current models cannot reliably handle real-world, iterative software engineering without human steering.

Author

humanlayer