# Preprint compares KV compression with extra GPUs for LLM serving

_Wednesday, August 26, 2026 at 12:00 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A preprint compared tensor parallelism with KV-cache compression for memory-bound LLM serving using a simulator calibrated on A100, A40 and H100 hardware. Across Llama-2 7B and 70B configurations, the authors report compression was 1.20x to 2.00x cheaper for the memory relief they modeled. They found tensor parallelism was necessary when model weights, rather than KV cache, exceeded a single GPU's memory, and reported that compression increased per-token latency by 8% to 93%.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.23962)

---
Canonical: https://techandbusiness.org/newswire/g1BkJYgtmGNaiY603vDhGm
Retrieved: 2026-08-26T14:33:55.491Z
Publisher: Tech & Business (techandbusiness.org)
