# Preprint proposes lazy memory retrieval for long LLM chats

_Tuesday, September 15, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A preprint describes Pull, a session router intended to reduce the context supplied to stateful LLM conversations. It maintains a local metadata directory and retrieves only turns needed for a query, while leaving collapsed turns available for later expansion. On the reported LoCoEval tests, Pull reduced per-query context tokens by 75.1% on single-hop tasks and 72.0% on multi-hop tasks with no reported quality loss. The results are from controlled benchmarks, not a production deployment.

## Sources

- [cs.CL updates on arXiv.org](https://arxiv.org/abs/2609.14773)

---
Canonical: https://techandbusiness.org/newswire/gaWnljXmj2WkR3lH1xGSRv
Retrieved: 2026-09-15T16:07:01.192Z
Publisher: Tech & Business (techandbusiness.org)
