Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Monea, Giovanni, Feldman, Yair, Padmanabhan, Shankar, Brantley, Kianté, Artzi, Yoav
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917172978122752
author Monea, Giovanni
Feldman, Yair
Padmanabhan, Shankar
Brantley, Kianté
Artzi, Yoav
author_facet Monea, Giovanni
Feldman, Yair
Padmanabhan, Shankar
Brantley, Kianté
Artzi, Yoav
contents The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit that as a model generates reasoning tokens, the informational value of past generated tokens diminishes, creating an opportunity for compression. In this work, we propose to periodically compress the generation KV cache with a learned, special-purpose token and evict compressed entries. We train the model to perform this compression via a modified joint distillation and reinforcement learning (RL) framework. Our training method minimizes overhead over the conventional RL process, as it leverages RL outputs for distillation. Empirically, our method achieves a superior memory-accuracy Pareto frontier compared to both the model without cache compression and training-free compression techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
Monea, Giovanni
Feldman, Yair
Padmanabhan, Shankar
Brantley, Kianté
Artzi, Yoav
Computation and Language
The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit that as a model generates reasoning tokens, the informational value of past generated tokens diminishes, creating an opportunity for compression. In this work, we propose to periodically compress the generation KV cache with a learned, special-purpose token and evict compressed entries. We train the model to perform this compression via a modified joint distillation and reinforcement learning (RL) framework. Our training method minimizes overhead over the conventional RL process, as it leverages RL outputs for distillation. Empirically, our method achieves a superior memory-accuracy Pareto frontier compared to both the model without cache compression and training-free compression techniques.
title Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
topic Computation and Language
url https://arxiv.org/abs/2510.13797