# Preprint reports DiffusionGemma reaches 1,500 tokens per second on one H100

_Published Thursday, August 20, 2026 at 1:18 PM EDT · Science, AI · Latest · Tier 2 — Notable_

A technical report on arXiv introduces DiffusionGemma, an experimental open-weight language model that generates text through discrete diffusion rather than token-by-token decoding.

The authors say it refines 256-token blocks in parallel and averages about 20 tokens per forward pass. They report roughly 1,500 output tokens per second on a single NVIDIA H100 across their evaluation suite. The model was fine-tuned from a 25.2-billion-parameter mixture-of-experts Gemma 4 model using a two-stage pipeline that the report says used fewer than 10% of the starting model's training-token budget.

## Sources

- [arxiv.org](https://arxiv.org/abs/2608.00146)

---
Canonical: https://techandbusiness.org/newswire/D61nysEcdDwp8aBO4s2-PQ
Published: 2026-08-20T17:18:16.673Z
Story chronology: 2026-08-20T13:24:32.000Z
Retrieved: 2026-10-04T21:05:57.113Z
Publisher: Tech & Business (techandbusiness.org)
