
SWE-Bench Pro, a benchmark designed to improve on SWE-bench Verified, was found to have significant issues, with ~30% of its tasks broken. The problems include overly strict tests, underspecified prompts, low-coverage checks, and misleading prompts, affecting model reliability and safety assessments. A human-supervised review found more broken tasks than an automated filter, highlighting the need for more thorough data quality checks. The audit suggests that the benchmark is not easily trusted and emphasizes the importance of creating hard, fair benchmarks to evaluate model capabilities.
See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.

OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
