Anthropic Releases Automated Alignment Researchers for Reproducible AI Safety Research
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Anthropic's AARs represent a major step in automated AI safety research.
Anthropic released Automated Alignment Researchers (AARs), a Claude-powered research sandbox using nine Claude Opus 4.6 agents in independent sandboxes with a shared forum and codebase to automate alignment experiments. On a chat-task benchmark, the AARs achieved a performance gap recovered (PGR) of ~0.97 after ~800 cumulative AAR hours at ~$18,000 compute cost, far exceeding a human baseline PGR of 0.23 over seven days. However, results were uneven across math (PGR 0.94) and coding (PGR 0.47) tasks, underscoring that strong benchmark performance does not guarantee reliability across all scenarios.