# TaSQ preprint reports higher AI throughput with 1-bit cache compression

_Published Monday, October 5, 2026 at 6:04 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in a preprint that TaSQ, a method for compressing the memory used during language-model inference, supports up to 14× larger batch sizes and achieves 1.87× higher peak throughput than a BF16 baseline on a single RTX 6000 Ada GPU. The results come from an implementation in SGLang.

TaSQ compresses the key-value cache, which stores information used to process subsequent tokens. It weights and groups cached channels to reduce compression errors, with transforms that can be merged into model weights and compression codebooks. The researchers report better results than existing low-bit baselines across general, reasoning and long-context retrieval benchmarks; the stated throughput comparison is specific to that GPU implementation.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.03027)

---
Canonical: https://techandbusiness.org/newswire/GL7GtbWvx3kZeMlegx_jji
Published: 2026-10-05T10:04:14.061Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T11:30:55.671Z
Publisher: Tech & Business (techandbusiness.org)
