Skip to main content

Share story

AI Science

SpAx preprint reports faster decoding for models stored outside GPU memory

Researchers report in a preprint that SpAx accelerates language-model decoding when model weights reside in CPU memory. Average speedups are 3.86X for 16-bit weights and 2.06X for 4-bit weights, with a WikiText-2 perplexity increase of at most 10%, indicating a loss in predictive quality. SpAx reduces data transfers by skipping weights associated with activations closest to zero, transferring compressed approximations for smaller activations and retaining original weights for the largest activations. For weights stored in flash storage, average speedups are 3.31X and 1.54X, respectively. The method addresses transfer bottlenecks on consumer GPUs that cannot hold all model weights.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.LG updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Pinnacle acquires Qmulus Solutions' Sage customer base

Pinnacle has acquired the Sage customer base of Evesham-based Qmulus Solutions, transferring more than 100 customers to the technology solutions provider. The deal covers customers using Sage 200, Sage Intacct and Sage CRM. Custo...

Infrastructure
Infrastructure

VSMC opens Singapore wafer fab and enters risk production

VisionPower Semiconductor Manufacturing Company has opened its first 300mm wafer fab in Tampines, Singapore, and entered risk production, an initial manufacturing phase ahead of commercial volume output. The joint venture between ...

Capital
Capital

IRACE Digital Bank acquires blockchain startup Trrue

IRACE Digital Bank, formerly Fundbank, has acquired Irish blockchain startup Trrue in a transaction valued at $11.8 million, CoinTrust reports. Dealroom reported a different value of $11.2 million. The acquisition brings Trrue's ...