{"version":"1.0","type":"rich","provider_name":"techandbusiness.org","provider_url":"https://techandbusiness.org","title":"vLLM 0.28 adds KV offloading and serving changes","author_name":"techandbusiness.org · Infrastructure","thumbnail_url":"https://opengraph.githubassets.com/4ebc6bdc52e9be13e062849c7d527d88f083515ccff086b45b804e1e708e9324/vllm-project/vllm/releases/tag/v0.28.0","width":600,"height":400,"html":"<blockquote class=\"tb-newswire-embed\" style=\"max-width:600px;border-left:3px solid #22d3ee;padding:12px 16px;margin:0;font-family:-apple-system,system-ui,sans-serif;background:#09090b;border-radius:0 8px 8px 0;\">\n      <p style=\"margin:0 0 8px;font-size:10px;font-weight:600;letter-spacing:0.1em;color:#71717a;\">techandbusiness.org · Infrastructure</p>\n      <p style=\"margin:0 0 8px;font-size:18px;font-weight:700;line-height:1.3;color:#fff;\"><a href=\"https://techandbusiness.org/newswire/Qm_v3PIK7C_g-kVhIq8Pub\" style=\"color:#fff;text-decoration:none;\">vLLM 0.28 adds KV offloading and serving changes</a></p>\n      <p style=\"margin:0;font-size:14px;color:#a1a1aa;line-height:1.5;\">vLLM released version 0.28.0, adding disk-based KV-cache offloading, E/P/D disaggregation in Model Runner V2, and a Rust frontend with gRPC multimodal image inference. The project also raised the defa...</p>\n    </blockquote>"}