Skip to main content

Share story

AI Science

TaSQ preprint reports higher AI throughput with 1-bit cache compression

Researchers report in a preprint that TaSQ, a method for compressing the memory used during language-model inference, supports up to 14× larger batch sizes and achieves 1.87× higher peak throughput than a BF16 baseline on a single RTX 6000 Ada GPU. The results come from an implementation in SGLang. TaSQ compresses the key-value cache, which stores information used to process subsequent tokens. It weights and groups cached channels to reduce compression errors, with transforms that can be merged into model weights and compression codebooks. The researchers report better results than existing low-bit baselines across general, reasoning and long-context retrieval benchmarks; the stated throughput comparison is specific to that GPU implementation.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.LG updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Pinnacle acquires Qmulus Solutions' Sage customer base

Pinnacle has acquired the Sage customer base of Evesham-based Qmulus Solutions, transferring more than 100 customers to the technology solutions provider. The deal covers customers using Sage 200, Sage Intacct and Sage CRM. Custo...

Infrastructure
Infrastructure

VSMC opens Singapore wafer fab and enters risk production

VisionPower Semiconductor Manufacturing Company has opened its first 300mm wafer fab in Tampines, Singapore, and entered risk production, an initial manufacturing phase ahead of commercial volume output. The joint venture between ...

Capital
Capital

IRACE Digital Bank acquires blockchain startup Trrue

IRACE Digital Bank, formerly Fundbank, has acquired Irish blockchain startup Trrue in a transaction valued at $11.8 million, CoinTrust reports. Dealroom reported a different value of $11.2 million. The acquisition brings Trrue's ...