# Preprint tests uncertainty estimates for closed language models

_Published Wednesday, September 23, 2026 at 5:08 PM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers introduced Pinocchio, a tool that estimates whether a language model's answer is correct without access to the model's internal data. In a preprint, they report that a version trained on responses from seven models scored 0.862 on a measure of how well it distinguished correct from incorrect answers held out from those models.

The tool needs one processing pass and also transferred to thirteen previously unseen models across eight organizations. The researchers released code, but the reported score comes from models represented in its training data.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.24881)

---
Canonical: https://techandbusiness.org/newswire/PBFk528tB-KTuZqumDE9da
Published: 2026-09-23T21:08:18.101Z
Story chronology: 2026-09-23T04:00:00.000Z
Retrieved: 2026-09-23T22:32:26.446Z
Publisher: Tech & Business (techandbusiness.org)
