SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Abstract
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.
Community
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some frontier models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark version for assessing software engineering agents.
Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.
Figure 1 is exactly the comparison between the original score and the verified score.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring (2026)
- Shortcutting the Fix: Identifying and Categorizing Agentic Exploits in Software Engineering Benchmarks (2026)
- SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents (2026)
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (2026)
- BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP (2026)
- ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders (2026)
- OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.08149 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper