# Self-play training study reports gains in mathematical reasoning

_Published Monday, September 28, 2026 at 1:55 PM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report that training a Qwen3-4B-Base model on records from self-play board-game searches raised its mean score across six mathematics benchmarks from 24.1 to 36.6. The model's win rate on a held-out game also rose from 15% to 45%.

Their preprint describes a method that turns game searches into training examples containing a preferred move, alternatives, possible replies and estimated outcomes. The reported transfer to unseen mathematics suggests a way to generate reasoning training data without human labeling. The results come from the tested model and benchmarks; the source does not establish the same gains for other models or workloads.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.30936)

---
Canonical: https://techandbusiness.org/newswire/5TcUf_VgBEsW0UqEbowrB8
Published: 2026-09-28T17:55:23.014Z
Story chronology: 2026-09-28T04:00:00.000Z
Retrieved: 2026-09-28T19:55:07.655Z
Publisher: Tech & Business (techandbusiness.org)
