# FP4 pretraining study reports faster throughput with training-loss tradeoffs

_Published Saturday, October 3, 2026 at 12:12 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in an arXiv preprint that a custom four-bit floating-point training route reached 37.9K tokens/s/GPU, compared with 18.8K for bfloat16 and 27.6K for Transformer Engine in matched tests on the same accelerator. They evaluated Llama-3-family 8B pretraining through 160 billion tokens.

Their method coordinates quantization-the conversion of values into a lower-precision format-with scaling and the data layouts consumed by subsequent operations, addressing overhead that can erase faster matrix multiplication.

Another route reached 37.2K tokens/s/GPU but finished with training loss 2.11% above the raw bfloat16 endpoint. Rankings on downstream tasks differed from training-loss rankings, showing that the faster execution paths do not provide a uniform quality advantage.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.00053)

---
Canonical: https://techandbusiness.org/newswire/uVlhnD2VPzKxSTdPOttlIX
Published: 2026-10-03T04:12:54.430Z
Story chronology: 2026-10-03T04:00:00.000Z
Retrieved: 2026-10-03T06:37:01.739Z
Publisher: Tech & Business (techandbusiness.org)
