# Preprint reports DiffusionGemma reaches 1,500 tokens per second on one H100

_Thursday, August 20, 2026 at 9:24 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A technical report on arXiv introduces DiffusionGemma, an experimental open-weight language model that generates text through discrete diffusion rather than token-by-token decoding.

The authors say it refines 256-token blocks in parallel and averages about 20 tokens per forward pass. They report roughly 1,500 output tokens per second on a single NVIDIA H100 across their evaluation suite. The model was fine-tuned from a 25.2-billion-parameter mixture-of-experts Gemma 4 model using a two-stage pipeline that the report says used fewer than 10% of the starting model's training-token budget.

## Sources

- [arxiv.org](https://arxiv.org/abs/2608.00146)

---
Canonical: https://techandbusiness.org/newswire/D61nysEcdDwp8aBO4s2-PQ
Retrieved: 2026-08-20T19:29:35.016Z
Publisher: Tech & Business (techandbusiness.org)
