Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Sangyun, McLeish, Sean, Goldstein, Tom, Fanti, Giulia
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916055330324480
author Lee, Sangyun
McLeish, Sean
Goldstein, Tom
Fanti, Giulia
author_facet Lee, Sangyun
McLeish, Sean
Goldstein, Tom
Fanti, Giulia
contents Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache. During sleep, the model performs $N$ offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule. During inference, this shifts extra computation to sleep while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration $N$ for our models improves performance, with the largest gains on examples that require deeper reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26099
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
Lee, Sangyun
McLeish, Sean
Goldstein, Tom
Fanti, Giulia
Computation and Language
Artificial Intelligence
Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache. During sleep, the model performs $N$ offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule. During inference, this shifts extra computation to sleep while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration $N$ for our models improves performance, with the largest gains on examples that require deeper reasoning.
title Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.26099