Skip to main content

Share story

AI Science

Preprint reports faster GB200 training with polynomial function replacements

Researchers report in an arXiv preprint that replacing selected mathematical functions with short polynomial calculations improved complete training-step throughput on GB200 hardware by 2.7%, 2.9%, and 8.0% across three tasks. The replacements approximate functions used inside language models, combining symmetry, rounding to the target number format, and packed arithmetic within the kernels that consume their outputs. A fourth task, sigmoid attention, improved complete-attention forward performance by 7.4% but the complete GPU step by 0.3%. The evaluation included one paired pretraining comparison per task. At common horizons near 100 billion tokens, final smoothed training-loss differences between polynomial and native implementations ranged from -0.107 to +0.079, showing that the substitutions also changed model training behavior.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.LG updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Science
AI Science

Trillium Labs launches with plans to publish AI training experiments

Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit that plans to publish AI experiment details so outside researchers can study and replicate its work, WIRED reported. The lab has raised an undisclosed sum from Schmid...

Science AI
Science AI

Lean plans four proof-checking kernels to guard against AI exploits

Lean's next version will ship with four different proof-checking kernels instead of one, New Scientist reported. These components check the logic of mathematical proofs, and the change follows an AI-assisted stunt that exploited s...