Skip to main content

Share story

AI Science

CARGO study improves AI agent answer checks but misses procedural errors

Researchers testing a proposed method for evaluating AI agents found that it avoided false penalties when a correct answer used facts from a live case rather than identifiers from a reference case. In a 246-item diagnostic benchmark, standard reference-based judges penalized all 50 correct answers transferred to another case; the CARGO method penalized none while detecting 50 of 50 and 49 of 50 contradictory answers across two judge models. CARGO checks claims against the live case and withholds judgment when retrieval confidence is low. The preprint also found that its more lenient approach detected only 20% of procedural errors, a gap a subsequent fix did not close.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Infrastructure
Infrastructure

Samsung commits $1 billion to KKR-backed AI infrastructure company Helix

Samsung has committed $1 billion to Helix Digital Infrastructure through a long-duration capital fund, Helix and KKR announced. The commitment adds to more than $10 billion previously committed to the Helix strategy by founding i...

AI Capital
AI Capital

Instinct confirms $1 billion Series C at $10 billion valuation

AI assistant startup Instinct confirmed a $1 billion Series C backed by investors including Sequoia Capital, Benchmark Capital and Coatue. The round values the company at $10 billion, up from the $2.5 billion valuation it announce...

Infrastructure Products
Infrastructure Products

Starship reaches orbit and deploys 26 Starlink satellites

SpaceX's Starship reached Earth orbit for the first time on Monday and deployed 26 new Starlink satellites. The flight continued after the upper stage lost one of its six engines shortly after separation. SpaceX later decided to b...