Skip to main content

Share story

AI Science

FTW preprint matches agent-training baselines without repeated group trials

Researchers report in a preprint that Follow the Winners, an algorithm for training language-model agents, matches GRPO and PPO on Sokoban and Search-R1 baselines while substituting stored samples in CPU memory for a value model or repeated groups of trials. The method ranks samples from a replay buffer by their returns and filters them to guide training. It targets environments such as live services and security sandboxes, where repeating agent trajectories can be impractical. The researchers identify a bounded preference for riskier outcomes as a tradeoff of this filtering approach, which FTW controls.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.LG updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Pinnacle acquires Qmulus Solutions' Sage customer base

Pinnacle has acquired the Sage customer base of Evesham-based Qmulus Solutions, transferring more than 100 customers to the technology solutions provider. The deal covers customers using Sage 200, Sage Intacct and Sage CRM. Custo...

Infrastructure
Infrastructure

VSMC opens Singapore wafer fab and enters risk production

VisionPower Semiconductor Manufacturing Company has opened its first 300mm wafer fab in Tampines, Singapore, and entered risk production, an initial manufacturing phase ahead of commercial volume output. The joint venture between ...

Capital
Capital

IRACE Digital Bank acquires blockchain startup Trrue

IRACE Digital Bank, formerly Fundbank, has acquired Irish blockchain startup Trrue in a transaction valued at $11.8 million, CoinTrust reports. Dealroom reported a different value of $11.2 million. The acquisition brings Trrue's ...