# Preprint reports task-completion gains from agent-generated interfaces

_Published Friday, October 9, 2026 at 12:05 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves information and executes tasks with another that builds interfaces to resolve ambiguities and structure user input.

The researchers introduced a benchmark built on 10 domain databases, with 300-task and 1,000-task splits. On the smaller split, the harness gained an average 4.48 percentage points over smolagents. A reviewer survey found that generated interfaces reduced average dialogue rounds from 3.4 to 1.2; the findings concern evaluated database-backed workflows.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2610.11123)

---
Canonical: https://techandbusiness.org/newswire/XagL5akZRf_hZWRMs8FZD3
Published: 2026-10-09T04:05:07.468Z
Story chronology: 2026-10-09T04:00:00.000Z
Retrieved: 2026-10-09T06:33:01.306Z
Publisher: Tech & Business (techandbusiness.org)
