KaVa: Latent Reasoning via Compressed KV-Cache Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuzina, Anna, Pioro, Maciej, Whatmough, Paul N., Bejnordi, Babak Ehteshami
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909020728590336
author Kuzina, Anna
Pioro, Maciej
Whatmough, Paul N.
Bejnordi, Babak Ehteshami
author_facet Kuzina, Anna
Pioro, Maciej
Whatmough, Paul N.
Bejnordi, Babak Ehteshami
contents Large Language Models (LLMs) excel at multi-step reasoning problems with explicit chain-of-thought (CoT), but verbose traces incur significant computational costs and memory overhead, and often carry redundant, stylistic artifacts. Latent reasoning has emerged as an efficient alternative that internalizes the thought process, but it suffers from a critical lack of supervision, limiting its effectiveness on complex, natural-language reasoning traces. In this work we propose KaVa, the first framework that bridges this gap by distilling knowledge directly from a compressed KV-cache of the teacher into a latent-reasoning student via self-distillation, leveraging the representational flexibility of continuous latent tokens to align stepwise KV trajectories. We show that the abstract, unstructured knowledge within compressed KV-cache, which lacks direct token correspondence, can serve as a rich supervisory signal for a latent reasoning student. Empirically, the approach consistently outperforms strong latent baselines, exhibits markedly smaller degradation from equation-only to natural-language traces, and scales to larger backbones while preserving efficiency. These results establish compressed KV-cache distillation as a scalable supervision signal for latent reasoning, combining the accuracy of CoT-trained teachers with the efficiency and deployability of latent inference.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KaVa: Latent Reasoning via Compressed KV-Cache Distillation
Kuzina, Anna
Pioro, Maciej
Whatmough, Paul N.
Bejnordi, Babak Ehteshami
Machine Learning
Large Language Models (LLMs) excel at multi-step reasoning problems with explicit chain-of-thought (CoT), but verbose traces incur significant computational costs and memory overhead, and often carry redundant, stylistic artifacts. Latent reasoning has emerged as an efficient alternative that internalizes the thought process, but it suffers from a critical lack of supervision, limiting its effectiveness on complex, natural-language reasoning traces. In this work we propose KaVa, the first framework that bridges this gap by distilling knowledge directly from a compressed KV-cache of the teacher into a latent-reasoning student via self-distillation, leveraging the representational flexibility of continuous latent tokens to align stepwise KV trajectories. We show that the abstract, unstructured knowledge within compressed KV-cache, which lacks direct token correspondence, can serve as a rich supervisory signal for a latent reasoning student. Empirically, the approach consistently outperforms strong latent baselines, exhibits markedly smaller degradation from equation-only to natural-language traces, and scales to larger backbones while preserving efficiency. These results establish compressed KV-cache distillation as a scalable supervision signal for latent reasoning, combining the accuracy of CoT-trained teachers with the efficiency and deployability of latent inference.
title KaVa: Latent Reasoning via Compressed KV-Cache Distillation
topic Machine Learning
url https://arxiv.org/abs/2510.02312