# Preprint reports 80% fewer evaluations for LLM-agent harness optimization

_Friday, August 21, 2026 at 12:00 AM EDT · Science, AI · Latest · Tier 2 — Notable_

An arXiv preprint presents Task-CoEvolve, a method for optimizing LLM-agent harnesses by changing which validation tasks are sampled as the harness evolves. The authors say experiments on online text classification and Terminal-Bench 2.1 matched the final performance of full-set search while reducing optimization evaluations by 80% versus that approach. The method samples tasks where candidate harnesses disagree and estimates full-set scores from those partial evaluations. The work is a preprint, and the authors say code will be released.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.20169)

---
Canonical: https://techandbusiness.org/newswire/vQO2DoKWix3h1vKlp1ub_j
Retrieved: 2026-08-21T10:08:32.380Z
Publisher: Tech & Business (techandbusiness.org)
