AI Science
Preprint finds single-run agent audits can miss irreversible damage
A preprint introducing AgentRelBench reports that irreversible database-state damage occurred across the measured model families but not on every run, making one-shot audits unreliable for detecting some harmful agent behavior. Across 2,128 runs, the authors found a clean run missed a damage-producing model-task pair 0.80 of the time in the development pool. The held-out result was described as underpowered, and the capability gradient was observational rather than causal.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire