# QF3 researchers report 10x faster training for humanoid robot policies

_Published Tuesday, October 6, 2026 at 11:07 PM EDT · Science, Robotics · Latest · Tier 2 — Notable_

Researchers introduce QF3, a reinforcement learning algorithm that they report trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++. Their preprint also reports training humanoid locomotion policies from scratch and transferring them to hardware without additional training.

QF3 combines flow matching with feedback from a critic that evaluates actions. It restricts that feedback to action dimensions close to previously recorded actions, where the estimates are more reliable. The speed result pairs the algorithm with a high-throughput training recipe. Researchers also tested it for refining manipulation policies learned from demonstrations on ABC-Sim and Robomimic tasks.

## Sources

- [arXiv Query: search_query=cat:cs.RO&id_list=&start=0&max_results=30](https://arxiv.org/abs/2610.08789v1)

---
Canonical: https://techandbusiness.org/newswire/3c6NaNvJ7RTtvYZBa0R7Yp
Published: 2026-10-07T03:07:26.369Z
Story chronology: 2026-10-06T17:59:34.000Z
Retrieved: 2026-10-07T04:45:59.776Z
Publisher: Tech & Business (techandbusiness.org)
