Skip to content

How a Strands agent took Claude Opus 5 from 30% to 99.95% on ARC-AGI-3

7.8 relevance
Score Breakdown
technical depth
9
novelty
9
actionability
5
community
6
strategic
8
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Breakthrough in AI agent reasoning on ARC-AGI benchmark with deep technical detail.

AI/ML dev.to
How a Strands agent took Claude Opus 5 from 30% to 99.95% on ARC-AGI-3
Summary

AWS engineers used the open-source Strands Agents SDK with Claude Opus 5 to achieve a 99.95% relative human action efficiency (RHAE) score on ARC-AGI-3's public benchmark, completing all 183 levels across 25 environments in a single 8-hour run costing $830 in tokens. The agent started with zero game knowledge in its system prompt, using generic action labels and raw numeric board representations, forcing it to experiment and learn through observation. This contrasts sharply with Claude Opus 5's standalone score of 30.16% on the same benchmark, demonstrating that agent harness design—not just model capability—determines performance on long-horizon novel problem-solving tasks.

Author

Morgan Willis

More from Morgan Willis →