# SpAx preprint reports faster decoding for models stored outside GPU memory

_Published Monday, October 5, 2026 at 5:57 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in a preprint that SpAx accelerates language-model decoding when model weights reside in CPU memory. Average speedups are 3.86X for 16-bit weights and 2.06X for 4-bit weights, with a WikiText-2 perplexity increase of at most 10%, indicating a loss in predictive quality.

SpAx reduces data transfers by skipping weights associated with activations closest to zero, transferring compressed approximations for smaller activations and retaining original weights for the largest activations. For weights stored in flash storage, average speedups are 3.31X and 1.54X, respectively. The method addresses transfer bottlenecks on consumer GPUs that cannot hold all model weights.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.02598)

---
Canonical: https://techandbusiness.org/newswire/FD9EzeDzHAsLlsm0VsqBe6
Published: 2026-10-05T09:57:27.736Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T11:31:26.084Z
Publisher: Tech & Business (techandbusiness.org)
