Skip to content

Most coding agent benchmarks skip large-scale refactoring. Not this one.

7.4 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
7
community
5
strategic
6
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Benchmark for large-scale refactoring is highly actionable and directly relevant to AI coding agents.

AI/ML thenewstack.io
Most coding agent benchmarks skip large-scale refactoring. Not this one.
Summary

A new benchmark, SWE-Bench ProMax, targets large-scale code refactoring across seven languages (Python, Java, TypeScript, Go, C, C++, Rust) with 170 curated instances from real commits, finding the best AI agent achieves only a 41.2% resolve rate. Researchers at Shanghai Jiao Tong University, Peking University, and Douyin Group designed it to address quality flaws in existing benchmarks like SWE-Bench Verified, where nearly 60% of unsolved instances had flawed tests. Experts note LLMs lack the structural understanding needed for zero-tolerance refactoring, as token proximity does not guarantee comprehension of large codebases.

Author

Meredith Shubel

More from Meredith Shubel →