# Preprint finds commit-first LLM judges can inherit model errors

_Wednesday, September 2, 2026 at 12:00 AM EDT · Science · Latest · Tier 2 — Notable_

A preprint auditing 24 default LLM-judge configurations across eight evaluation frameworks found none used commit-first judging. In controlled tests, a best-of-N search produced candidates accepted by a documented judge configuration despite failing held-out tests; commit-first removed that effect on one task but worsened results on another when the judge's committed answer was wrong. The authors also report errors in five of their own fifteen source-based claims.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.00088)

---
Canonical: https://techandbusiness.org/newswire/uS79b3LFgmbGdokm86sipz
Retrieved: 2026-09-02T22:13:16.976Z
Publisher: Tech & Business (techandbusiness.org)
