# vLLM 0.28 adds KV offloading and serving changes

_Saturday, August 29, 2026 at 2:22 PM EDT · Infrastructure · Latest · Tier 2 — Notable_

![vLLM 0.28 adds KV offloading and serving changes — Primary](https://opengraph.githubassets.com/4ebc6bdc52e9be13e062849c7d527d88f083515ccff086b45b804e1e708e9324/vllm-project/vllm/releases/tag/v0.28.0)

vLLM released version 0.28.0, adding disk-based KV-cache offloading, E/P/D disaggregation in Model Runner V2, and a Rust frontend with gRPC multimodal image inference. The project also raised the default maximum batched tokens from 8,192 to 16,384 and enabled prefix caching by default for Mamba models. It lists release wheels for CUDA, CPU and other targets, while moving bitsandbytes support to an out-of-tree plugin and removing several deprecated options.

## Sources

- [GitHub](https://github.com/vllm-project/vllm/releases/tag/v0.28.0)

---
Canonical: https://techandbusiness.org/newswire/Qm_v3PIK7C_g-kVhIq8Pub
Retrieved: 2026-08-29T23:57:27.842Z
Publisher: Tech & Business (techandbusiness.org)
