Skip to main content
Back to Newswire
AI

DeepSeek Releases DSpark: Speculative Decoding Makes V4 Up to 85 Percent Faster

DeepSeek Image: Primary
DeepSeek on June 27 released DSpark, an inference optimization framework using speculative decoding that the company says makes its V4-Flash model generate responses up to 85 percent faster than the prior single-token baseline. The speed gain comes without retraining the model, changing its weights, or adding new hardware, according to DeepSeek. The framework is now live across V4-Flash and V4-Pro, and is available as open-source code under an MIT license. DeepSeek also released DeepSpec, a full-stack codebase for training and evaluating speculative decoding draft models, under an MIT license on GitHub. DeepSpec targets the Qwen3 and Gemma model families. The deployed configuration, called DSpark-5, uses a five-token draft block. In DeepSeek's internal production data, DSpark-5 improved per-user generation speed by 60 to 85 percent on V4-Flash and 57 to 78 percent on V4-Pro compared to the prior MTP-1 baseline. DeepSeek emphasized that DSpark is not a new model -- the Hugging Face cards for DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark use the same checkpoint with a speculative decoding module attached. No independent third-party verification of the claims has been published as of June 28, 2026.
Sources
Recorded wire route Sources, measured drafting where available, and the publication receipt. See concurrent Machine
Evidence entered
Admission Evidence and chronology passed AI
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from TechTimes, TechStartups and reviewed by the T&B editorial agent team.
Back to Newswire