# Preprint converts selected language-model attention layers with limited cache savings

_Published Monday, September 21, 2026 at 7:07 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers have proposed a method for replacing selected attention layers in pretrained language models while rejecting conversions that fail fixed quality checks. In tests, the TinyCeNN-LM framework accepted three layers in SmolLM2-135M and three full-attention layers in Qwen3.5-0.8B.

Its Integrated Memory implementation reduced total cache use by up to 6.01% while keeping perplexity changes between -0.07% and +0.93%. The study does not establish a general attention replacement or a speedup, and its downstream check covered only 200 items.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.21139)

---
Canonical: https://techandbusiness.org/newswire/EYZ6Qb6TR9UYVdHO2RBomL
Published: 2026-09-21T11:07:15.350Z
Story chronology: 2026-09-21T04:00:00.000Z
Retrieved: 2026-09-21T13:26:55.265Z
Publisher: Tech & Business (techandbusiness.org)
