Skip to main content
Back to Newswire
Science

Preprint finds LLM memory scores depend heavily on how history is rendered

A new arXiv preprint introduces RENDER, a benchmark control that holds a conversation fixed while changing how its history is presented to an answering model. Across 500 LongMemEval questions and nine models, the authors report that matched-budget resolved packets outperformed recency-truncated raw dialogue by 42.4 to 72.6 points. The study also found mixed model-specific significance after judge rescoring. The results are a preprint evaluation, not evidence of an operational memory-system deployment.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire