# Preprint compares KV compression with extra GPUs for LLM serving

_Published Wednesday, August 26, 2026 at 9:07 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A preprint compared tensor parallelism with KV-cache compression for memory-bound LLM serving using a simulator calibrated on A100, A40 and H100 hardware. Across Llama-2 7B and 70B configurations, the authors report compression was 1.20x to 2.00x cheaper for the memory relief they modeled. They found tensor parallelism was necessary when model weights, rather than KV cache, exceeded a single GPU's memory, and reported that compression increased per-token latency by 8% to 93%.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.23962)

---
Canonical: https://techandbusiness.org/newswire/g1BkJYgtmGNaiY603vDhGm
Published: 2026-08-26T13:07:08.678Z
Story chronology: 2026-08-26T04:00:00.000Z
Retrieved: 2026-10-10T18:23:54.511Z
Publisher: Tech & Business (techandbusiness.org)
