Modelwire
Subscribe

SWE-Bench ProMax resets code agent evaluation with refactoring focus

Existing code-generation benchmarks are collapsing under their own success. A new audit reveals that 60% of SWE-bench Verified tasks contain broken test cases, while frontier models exploit training-data leakage to artificially inflate scores. SWE-Bench ProMax addresses this crisis by shifting focus to multilingual code refactoring, a task class that demands coordinated, semantics-preserving edits across multiple files. This pivot matters because it forces agents to demonstrate genuine reasoning rather than pattern matching, resetting the evaluation bar for software engineering AI and exposing which systems have actually learned to code versus which have merely memorized.

Modelwire context

Skeptical read

SWE-Bench ProMax doesn't address why the original benchmark failed (training data leakage, broken test cases); it sidesteps the problem by moving to a new task domain. That's a valid triage move, but it doesn't explain whether multilingual refactoring will resist the same contamination pressures that collapsed the prior benchmark, or if this is just buying time before the next audit finds similar rot.

This connects directly to the broader evaluation crisis surfaced in recent work on diagnostic stress testing (Decoding-Level Taboo, August) and token-level metric collapse (Mismatch Matters, same period). Those papers exposed how benchmarks measure performance under ideal conditions while missing fragility in real constraints. SWE-Bench ProMax is attempting to tighten the constraint set by forcing semantics-preserving edits across files, but it's still a benchmark-level fix, not a structural fix to how we validate that models have learned versus memorized. The Dutch municipal LLM evaluation framework from August offers a contrasting approach: instead of pivoting to harder tasks, it operationalizes multiple evaluation dimensions (factuality, bias, transparency) to catch different failure modes. SWE-Bench ProMax assumes one harder task will reveal truth; the Dutch work assumes truth requires multiple lenses.

If SWE-Bench ProMax scores remain stable across three consecutive model releases without evidence of training data leakage (check for exact task overlap in model cards and training documentation), the pivot worked. If a new audit within 12 months finds similar contamination or broken test cases in ProMax itself, the field has confirmed that task-switching is a temporary patch, not a solution.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsSWE-Bench · SWE-Bench ProMax · SWE-Bench Verified

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

SWE-Bench ProMax resets code agent evaluation with refactoring focus · Modelwire