# Preprint reports geometry-based preference tuning gains in LLM math reasoning

_Monday, August 17, 2026 at 12:00 AM EDT · Science · Latest · Tier 2 — Notable_

A new arXiv preprint proposes Cloud-ScPO, a semi-supervised preference-optimization method that derives training pairs partly from the geometry of a model's hidden-state representations. The authors say the framework combines those signals with answer-level self-consistency to select and filter reasoning trajectories. Across GSM8K and MATH-Numeric in four model settings, the paper reports improvements over ScPO of up to 4.49% and 4.19%, respectively. The findings have not been independently verified.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.01014)

---
Canonical: https://techandbusiness.org/newswire/9OeJx2K_iFpKM0qKyhc2dl
Retrieved: 2026-08-17T11:50:21.000Z
Publisher: Tech & Business (techandbusiness.org)
