# Separating signal from noise in coding evaluations

_Published Thursday, July 9, 2026 at 4:53 PM EDT · AI · Latest · Tier 2 — Notable_

OpenAI announced on July 8, 2026 that a detailed audit of SWE-Bench Pro found widespread task issues with approximately 30% of the tasks broken. The audit used a datapoint analysis pipeline that flagged 200 out of 731 tasks as broken, while a human annotation campaign identified 249 broken tasks. Issues included overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advised model developers to carefully examine results from the benchmark.

## Sources

- [OpenAI](https://openai.com/index/separating-signal-from-noise-coding-evaluations)

---
Canonical: https://techandbusiness.org/newswire/1zexheHxKDYI99qGZrH1XO
Published: 2026-07-09T20:53:18.459Z
Story chronology: 2026-07-08T12:00:00.000Z
Retrieved: 2026-10-08T04:31:26.909Z
Publisher: Tech & Business (techandbusiness.org)
