# K-Bench preprint finds no tested scientific agent clears all judge thresholds

_Friday, September 4, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A new preprint reports K-Bench 01, an evaluation of nine frontier models on 1,602 completed scientific-agent runs drawn from first-turn requests sampled from K-Dense Web traffic. The authors say no model met their scientist-acceptance threshold under all three blinded language-model judges. They report gpt-5.6-sol had the highest pooled mean, while its confidence interval spanned the threshold and two judges ranked claude-opus-5 first. Across 39,934 scored judgments, 47.6% fell below the threshold.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.21601)

---
Canonical: https://techandbusiness.org/newswire/C4VBy5lzkFHjCtNk7c-QWo
Retrieved: 2026-09-04T10:27:51.802Z
Publisher: Tech & Business (techandbusiness.org)
