Skip to main content

Share story

AI

DeepSeek Releases DSpark: Speculative Decoding Makes V4 Up to 85 Percent Faster

DeepSeek Image: Primary
DeepSeek on June 27 released DSpark, an inference optimization framework using speculative decoding that the company says makes its V4-Flash model generate responses up to 85 percent faster than the prior single-token baseline. The speed gain comes without retraining the model, changing its weights, or adding new hardware, according to DeepSeek. The framework is now live across V4-Flash and V4-Pro, and is available as open-source code under an MIT license. DeepSeek also released DeepSpec, a full-stack codebase for training and evaluating speculative decoding draft models, under an MIT license on GitHub. DeepSpec targets the Qwen3 and Gemma model families. The deployed configuration, called DSpark-5, uses a five-token draft block. In DeepSeek's internal production data, DSpark-5 improved per-user generation speed by 60 to 85 percent on V4-Flash and 57 to 78 percent on V4-Pro compared to the prior MTP-1 baseline. DeepSeek emphasized that DSpark is not a new model -- the Hugging Face cards for DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark use the same checkpoint with a speculative decoding module attached. No independent third-party verification of the claims has been published as of June 28, 2026.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from TechTimes, TechStartups and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Robotics Science
Robotics Science

MIT team demonstrates thin muscle-powered swimming robot

MIT engineers demonstrated a thin robot powered by a single layer of living muscle cells that swam through a simple maze in a petri dish. The researchers grew light-responsive cells on two grooved gel fins. Flashing light on eithe...

AI Products
AI Products

OpenAI pauses training of its most capable models amid safety concerns

OpenAI has paused training of its most capable models following its disclosure of the Hugging Face breach and other incidents, The Next Web reported, citing Axios. An OpenAI spokesperson said training would resume once more safegu...

AI
AI

Jeff releases small decision models for local use

The Jeff project has released model weights and code for small AI models that choose among options described in plain language and return a probability for each choice. Its published models include 0.8B and 2B parameter versions b...