# F-GRPO preprint targets rare-correct trajectories in RL training

_Friday, September 4, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A preprint proposes F-GRPO, a training adjustment intended to reduce the chance that finite rollout groups overlook rare correct trajectories during reinforcement learning with verifiable rewards. The authors derive a model of those tail-miss events and apply a difficulty-aware coefficient that down-weights updates on high-success sampled groups. In experiments on Qwen2.5-7B with groups of eight, they report higher math pass@256 across GRPO, DAPO and CISPO without increasing group size or compute cost.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2602.06717)

---
Canonical: https://techandbusiness.org/newswire/aBNFgArX9q4q21671KHmU5
Retrieved: 2026-09-04T15:56:14.860Z
Publisher: Tech & Business (techandbusiness.org)
