# FTW preprint matches agent-training baselines without repeated group trials

_Published Monday, October 5, 2026 at 6:06 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in a preprint that Follow the Winners, an algorithm for training language-model agents, matches GRPO and PPO on Sokoban and Search-R1 baselines while substituting stored samples in CPU memory for a value model or repeated groups of trials.

The method ranks samples from a replay buffer by their returns and filters them to guide training. It targets environments such as live services and security sandboxes, where repeating agent trajectories can be impractical. The researchers identify a bounded preference for riskier outcomes as a tradeoff of this filtering approach, which FTW controls.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.03361)

---
Canonical: https://techandbusiness.org/newswire/lMYi-Ix0zT9y2xSQjIN2Pu
Published: 2026-10-05T10:06:15.683Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T11:32:21.405Z
Publisher: Tech & Business (techandbusiness.org)
