Science AI
Preprint reports 80% fewer evaluations for LLM-agent harness optimization
An arXiv preprint presents Task-CoEvolve, a method for optimizing LLM-agent harnesses by changing which validation tasks are sampled as the harness evolves. The authors say experiments on online text classification and Terminal-Bench 2.1 matched the final performance of full-set search while reducing optimization evaluations by 80% versus that approach. The method samples tasks where candidate harnesses disagree and estimates full-set scores from those partial evaluations. The work is a preprint, and the authors say code will be released.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire