# Preprint proposes evolving terminal-agent training environments

_Friday, September 4, 2026 at 12:00 AM EDT · Science · Latest · Tier 2 — Notable_

A preprint proposes "environment evolution," a method that incrementally increases terminal-agent environment difficulty off-policy during reinforcement-learning training. The authors say tests with three frontier models produced more difficult environments, then report that simple long-horizon RL training improved Qwen3.6-27B and Qwen3.6-35B-A3B by 14.4 and 18.0 percentage points, respectively, on Terminal-Bench 2.1. The source describes the tested models as including preview systems, and the results remain unverified.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.04128)

---
Canonical: https://techandbusiness.org/newswire/y_WCqhfPmjXqCLuhye03BZ
Retrieved: 2026-09-04T16:58:58.675Z
Publisher: Tech & Business (techandbusiness.org)
