# CARGO study improves AI agent answer checks but misses procedural errors

_Published Monday, September 28, 2026 at 10:08 PM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers testing a proposed method for evaluating AI agents found that it avoided false penalties when a correct answer used facts from a live case rather than identifiers from a reference case. In a 246-item diagnostic benchmark, standard reference-based judges penalized all 50 correct answers transferred to another case; the CARGO method penalized none while detecting 50 of 50 and 49 of 50 contradictory answers across two judge models.

CARGO checks claims against the live case and withholds judgment when retrieval confidence is low. The preprint also found that its more lenient approach detected only 20% of procedural errors, a gap a subsequent fix did not close.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.30471)

---
Canonical: https://techandbusiness.org/newswire/iB7kKIqoY501pl9OLkEsD3
Published: 2026-09-29T02:08:07.472Z
Story chronology: 2026-09-28T04:00:00.000Z
Retrieved: 2026-09-29T03:57:10.098Z
Publisher: Tech & Business (techandbusiness.org)
