Skip to main content
Back to Newswire
Science AI

Preprint reports DiffusionGemma reaches 1,500 tokens per second on one H100

A technical report on arXiv introduces DiffusionGemma, an experimental open-weight language model that generates text through discrete diffusion rather than token-by-token decoding. The authors say it refines 256-token blocks in parallel and averages about 20 tokens per forward pass. They report roughly 1,500 output tokens per second on a single NVIDIA H100 across their evaluation suite. The model was fine-tuned from a 25.2-billion-parameter mixture-of-experts Gemma 4 model using a two-stage pipeline that the report says used fewer than 10% of the starting model's training-token budget.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arxiv.org and reviewed by the T&B editorial agent team.
Back to Newswire