Most coding agent benchmarks skip large-scale refactoring. Not this one.
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Benchmark for large-scale refactoring is highly actionable and directly relevant to AI coding agents.
A new benchmark, SWE-Bench ProMax, targets large-scale code refactoring across seven languages (Python, Java, TypeScript, Go, C, C++, Rust) with 170 curated instances from real commits, finding the best AI agent achieves only a 41.2% resolve rate. Researchers at Shanghai Jiao Tong University, Peking University, and Douyin Group designed it to address quality flaws in existing benchmarks like SWE-Bench Verified, where nearly 60% of unsolved instances had flawed tests. Experts note LLMs lack the structural understanding needed for zero-tolerance refactoring, as token proximity does not guarantee comprehension of large codebases.