# Preprint finds shared model blocks cut GPU memory use by 18-48%

_Published Wednesday, September 23, 2026 at 11:07 AM EDT · AI, Science, Infrastructure · Latest · Tier 2 — Notable_

Researchers report that sharing parts of related language models cut GPU memory use by 18-48% in their tests. Their preprint describes LinkerLLM, a loader that lets variants derived from the same pretrained model share blocks of weights, the values that govern model behavior. The approach fit up to five 7B-parameter variants on one 24 GB consumer GPU.

Post-training changed every tensor the researchers examined, so identical-file matching saved no memory. Five of eight tested configurations retained at least 94% of the quality of an unshared variant on the reported benchmarks; each of the other three fell below that threshold on one benchmark. Independently trained specializations showed substantially less shared structure.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2609.26147)

---
Canonical: https://techandbusiness.org/newswire/VN1_sl06yAKc_b8NBZrq9s
Published: 2026-09-23T15:07:04.562Z
Story chronology: 2026-09-23T04:00:00.000Z
Retrieved: 2026-09-23T16:48:02.366Z
Publisher: Tech & Business (techandbusiness.org)
