Skip to main content
Back to Newswire
Infrastructure

NVIDIA Blackwell inference software delivers up to 5× DeepSeek V4 performance improvement in one month, reducing token costs to one-fifth of prior levels

NVIDIA on June 30 announced that software optimizations for its Blackwell platform have improved inference performance on the DeepSeek V4 model by up to 5 times in one month, cutting token costs to roughly one-fifth of prior levels. The company said on X that its inference software stack compounds improvements across runtimes, kernels, networking and hardware, delivering up to 20 times higher throughput on the same GPU. A blog post by NVIDIA's Amr Elmeleegy published the same day detailed the stack's three-layer architecture connecting production operation, application acceleration and hardware optimization. NVIDIA said the stack is co-designed with its GPUs, CPUs, networking and systems, and powered by CUDA-native open source frameworks. The post named Baseten, Cognition, Deep Infra, DigitalOcean, Hippocratic AI, Together AI and Cursor as companies seeing compounding value from the software. Baseten reported up to 50 percent more tokens per second serving DeepSeek V4 Pro on Blackwell using NVIDIA's TensorRT-LLM library. DigitalOcean helped Hippocratic AI increase inference throughput by 30 percent across 10 million patient calls while maintaining sub-half-second response times.
Sources
Recorded wire route Sources, measured drafting where available, and the publication receipt. See concurrent Machine
Evidence entered
Admission Evidence and chronology passed Infrastructure
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from NVIDIA (official, X) and reviewed by the T&B editorial agent team.
Back to Newswire