Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917172978122752 |
|---|---|
| author | Monea, Giovanni Feldman, Yair Padmanabhan, Shankar Brantley, Kianté Artzi, Yoav |
| author_facet | Monea, Giovanni Feldman, Yair Padmanabhan, Shankar Brantley, Kianté Artzi, Yoav |
| contents | The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit that as a model generates reasoning tokens, the informational value of past generated tokens diminishes, creating an opportunity for compression. In this work, we propose to periodically compress the generation KV cache with a learned, special-purpose token and evict compressed entries. We train the model to perform this compression via a modified joint distillation and reinforcement learning (RL) framework. Our training method minimizes overhead over the conventional RL process, as it leverages RL outputs for distillation. Empirically, our method achieves a superior memory-accuracy Pareto frontier compared to both the model without cache compression and training-free compression techniques. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13797 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons Monea, Giovanni Feldman, Yair Padmanabhan, Shankar Brantley, Kianté Artzi, Yoav Computation and Language The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit that as a model generates reasoning tokens, the informational value of past generated tokens diminishes, creating an opportunity for compression. In this work, we propose to periodically compress the generation KV cache with a learned, special-purpose token and evict compressed entries. We train the model to perform this compression via a modified joint distillation and reinforcement learning (RL) framework. Our training method minimizes overhead over the conventional RL process, as it leverages RL outputs for distillation. Empirically, our method achieves a superior memory-accuracy Pareto frontier compared to both the model without cache compression and training-free compression techniques. |
| title | Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.13797 |