AI
Separating signal from noise in coding evaluations
OpenAI announced on July 8, 2026 that a detailed audit of SWE-Bench Pro found widespread task issues with approximately 30% of the tasks broken. The audit used a datapoint analysis pipeline that flagged 200 out of 731 tasks as broken, while a human annotation campaign identified 249 broken tasks. Issues included overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advised model developers to carefully examine results from the benchmark.
Sources
Evidence entered
Admission Evidence and chronology passed AI
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from OpenAI and reviewed by the T&B editorial agent team.
Back to Newswire
