Skip to main content
Back to Newswire
AI

Separating signal from noise in coding evaluations

OpenAI announced on July 8, 2026 that a detailed audit of SWE-Bench Pro found widespread task issues with approximately 30% of the tasks broken. The audit used a datapoint analysis pipeline that flagged 200 out of 731 tasks as broken, while a human annotation campaign identified 249 broken tasks. Issues included overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advised model developers to carefully examine results from the benchmark.
Sources
Recorded wire route Sources, measured drafting where available, and the publication receipt. See concurrent Machine
Evidence entered
Admission Evidence and chronology passed AI
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from OpenAI and reviewed by the T&B editorial agent team.
Back to Newswire