# Preprint reports gains from training AI agents on generated edge cases

_Published Wednesday, September 23, 2026 at 7:15 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report that EdgeGen, a system for generating difficult tasks for AI agents, improved performance in an airline-service benchmark. It extracts compliance rules from an agent's specification and creates tasks grounded in a database that test those rules, without requiring people to label the tasks.

Fine-tuning with the generated data produced mean progress improvements of 2 percent to 42 percent on the tau2bench airline domain; some baseline methods reduced performance for some models. For one Gemma-4-e4b model, optimizing the agent's test setup improved mean progress by 10 percent over a human-curated setup and 30 percent over the base setup. The reported results are benchmark measurements.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.24115)

---
Canonical: https://techandbusiness.org/newswire/FQ512X2tnmMqykb7DL0Dbp
Published: 2026-09-23T11:15:35.322Z
Story chronology: 2026-09-23T04:00:00.000Z
Retrieved: 2026-09-23T14:59:02.731Z
Publisher: Tech & Business (techandbusiness.org)
