# SlimKV preprint reports faster decoding with compressed AI memory

_Published Monday, October 5, 2026 at 5:59 AM EDT · Science, AI · Latest · Tier 2 — Notable_

Researchers report in a preprint that SlimKV, a method for compressing the memory language models retain from earlier tokens, achieved up to 3.38x end-to-end decoding speedup over an uncompressed model at 128K context length.

The method trains models to store compressed context representations and perform attention directly on them, reducing the latency of reconstructing full-sized representations. On LongBench, it retained over 96% of the uncompressed model's score at 4x/8x compression and outperformed baselines at 16x/32x compression. The reported speedups are evaluation results, with attention alone accelerating by up to 7.34x.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.02953)

---
Canonical: https://techandbusiness.org/newswire/4QhUB5Fz5IeyQ5qVZ-zjFX
Published: 2026-10-05T09:59:24.654Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T11:32:48.794Z
Publisher: Tech & Business (techandbusiness.org)
