Skip to main content

Share story

AI Science

Preprint reports task-completion gains from agent-generated interfaces

Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves information and executes tasks with another that builds interfaces to resolve ambiguities and structure user input. The researchers introduced a benchmark built on 10 domain databases, with 300-task and 1,000-task splits. On the smaller split, the harness gained an average 4.48 percentage points over smolagents. A reviewer survey found that generated interfaces reduced average dialogue rounds from 3.4 to 1.2; the findings concern evaluated database-backed workflows.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Science
AI Science

Nullify preprint reports selective LLM forgetting without weight updates

Researchers report in a preprint that Nullify, a method designed to suppress specific memorized information in large language models, matches or surpasses established forgetting baselines on TOFU and MUSE while preserving model ut...

AI Science
AI Science

Preprint reports schema-free generation of valid enterprise test data

Researchers report in an arXiv preprint that their Generalist Populator agent generated enterprise data with 100% constraint satisfaction and 0.88 average marginal fidelity across ten simulated environments without accessing datab...

Robotics AI
Robotics AI

DreamTrue researchers report fewer interaction defects in robot video predictions

Researchers report that DreamTrue, a model that predicts videos of robot actions, reduced human-assessed interaction defects from 48.12% to 6.25% on AgiBot. The preprint addresses predictions that follow commands inaccurately or f...

Robotics AI
Robotics AI

FAITH preprint reports humanoid safety gains while preserving task performance

Researchers report that their FAITH safety filter achieved a 99.95% safety rate while retaining 97% of unfiltered task return on a 29-degree-of-freedom humanoid in Walking-Avoid. The preprint also describes demonstrations of the s...