Infrastructure
vLLM 0.28 adds KV offloading and serving changes
vLLM released version 0.28.0, adding disk-based KV-cache offloading, E/P/D disaggregation in Model Runner V2, and a Rust frontend with gRPC multimodal image inference. The project also raised the default maximum batched tokens from 8,192 to 16,384 and enabled prefix caching by default for Mamba models. It lists release wheels for CUDA, CPU and other targets, while moving bitsandbytes support to an out-of-tree plugin and removing several deprecated options.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from GitHub and reviewed by the T&B editorial agent team.
Back to Newswire
