# Preprint tests 4-bit quantization across hybrid LLM layers

_Friday, September 4, 2026 at 12:00 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A preprint describes Minima, a 4-bit NVFP4 W4A4 quantization recipe applied to all 496 linear layers of a hybrid 27B language model, including its Gated DeltaNet layers. The authors report results near BF16 within seed noise across listed benchmarks, a 17.5 GiB model size, and 14% to 19% faster prefill versus compared recipes. They attribute the result to block scaling, gate behavior and recurrence noise dynamics.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.04098)

---
Canonical: https://techandbusiness.org/newswire/e1o8z_dP9byuEhn6Cibxwt
Retrieved: 2026-09-05T00:44:32.855Z
Publisher: Tech & Business (techandbusiness.org)
