Skip to main content

Share story

Science AI

SlimKV preprint reports faster decoding with compressed AI memory

Researchers report in a preprint that SlimKV, a method for compressing the memory language models retain from earlier tokens, achieved up to 3.38x end-to-end decoding speedup over an uncompressed model at 128K context length. The method trains models to store compressed context representations and perform attention directly on them, reducing the latency of reconstructing full-sized representations. On LongBench, it retained over 96% of the uncompressed model's score at 4x/8x compression and outperformed baselines at 16x/32x compression. The reported speedups are evaluation results, with attention alone accelerating by up to 7.34x.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.LG updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Pinnacle acquires Qmulus Solutions' Sage customer base

Pinnacle has acquired the Sage customer base of Evesham-based Qmulus Solutions, transferring more than 100 customers to the technology solutions provider. The deal covers customers using Sage 200, Sage Intacct and Sage CRM. Custo...

Infrastructure
Infrastructure

VSMC opens Singapore wafer fab and enters risk production

VisionPower Semiconductor Manufacturing Company has opened its first 300mm wafer fab in Tampines, Singapore, and entered risk production, an initial manufacturing phase ahead of commercial volume output. The joint venture between ...

Capital
Capital

IRACE Digital Bank acquires blockchain startup Trrue

IRACE Digital Bank, formerly Fundbank, has acquired Irish blockchain startup Trrue in a transaction valued at $11.8 million, CoinTrust reports. Dealroom reported a different value of $11.2 million. The acquisition brings Trrue's ...