Science
Preprint reports geometry-based preference tuning gains in LLM math reasoning
A new arXiv preprint proposes Cloud-ScPO, a semi-supervised preference-optimization method that derives training pairs partly from the geometry of a model's hidden-state representations. The authors say the framework combines those signals with answer-level self-consistency to select and filter reasoning trajectories. Across GSM8K and MATH-Numeric in four model settings, the paper reports improvements over ScPO of up to 4.49% and 4.19%, respectively. The findings have not been independently verified.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire