0

SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard

https://labs.scale.com/blog/swe-bench-pro-v2(labs.scale.com)
A refreshed benchmark, SWE-Bench Pro V2, has been released to more accurately measure the real-world coding abilities of AI agents. This new version closes cheating loopholes by locking network access and re-grading on clean systems, with experts spending nearly 2,000 hours to fix flawed or ambiguous tasks. Initial results reveal a significant performance gap between the public benchmark and a private, unseen set, suggesting models often recall memorized solutions rather than genuinely problem-solve. Interestingly, studies also show that while different agent toolsets drastically change the cost and time spent, they have little impact on the final success rate for the most capable models.
0 pointsby hdt1 hour ago

Comments (0)

No comments yet. Be the first to comment!

Want to join the discussion?