Skip to content

Livenerf: Has Opus 5.5 been nerfed yet?

7.4 relevance
Score Breakdown
technical depth
7
novelty
7
actionability
8
community
9
strategic
5
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Open source tool for monitoring AI model behavior, directly actionable and highly relevant to the AI/ML workflow.

Open Source github.com
Benchmark for tracking model capability after release. - ninjahawk/livenerf
Summary

Livenerf is an open-source, deterministic benchmark built on the UK AISI's Inspect framework to detect post-launch degradation in frontier models like Claude Opus 5.5. Running daily via headless Claude Code, it measures statistical drift across a pre-registered panel of 78 GPQA/MMLU-Pro questions, finding that lower model effort reduces output tokens by 62% but accuracy by only 8.3 points. The benchmark currently cannot distinguish Opus 5.5 from Opus 5 at 99% confidence, highlighting the subtlety of stealth model changes.

Author

ninjahawk