I trust Claude for everything. This test made me rethink that.
6.9 relevance
Score Breakdown
technical depth 7
novelty 7
actionability 7
community 5
strategic 6
personal 9
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Claude vs Grok model comparison, useful for AI selection but not groundbreaking.
Summary
xAI's Grok 4.5 claims to match Anthropic's Claude Opus 4.8 on coding while using 4.2x fewer output tokens and costing significantly less ($2/$6 vs $5/$25 per million tokens). In a head-to-head test on three Rust tasks in the fd repo, both models produced identical bug fixes, but Grok used 4.3x fewer tokens. However, on the harder SWE-Bench Pro, Opus still outperforms Grok, so the claim is 'as good for far less' rather than superior.