Separating signal from noise in coding evaluations

Openai··Submitted by Mads Kristian Nylund
AI ToolsAI GovernanceAI Evaluation

SWE-Bench Pro, a benchmark designed to improve on SWE-bench Verified, was found to have significant issues, with ~30% of its tasks broken. The problems include overly strict tests, underspecified prompts, low-coverage checks, and misleading prompts, affecting model reliability and safety assessments. A human-supervised review found more broken tasks than an automated filter, highlighting the need for more thorough data quality checks. The audit suggests that the benchmark is not easily trusted and emphasizes the importance of creating hard, fair benchmarks to evaluate model capabilities.

Read Article

More from Openai

Related Articles