Latent Context Compilation: Distilling Long Context into Compact Portable Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zeju, Zhou, Yizhou, Xu, Qiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918355169968128
author Li, Zeju
Zhou, Yizhou
Xu, Qiang
author_facet Li, Zeju
Zhou, Yizhou
Xu, Qiang
contents Efficient long-context LLM deployment is stalled by a dichotomy between amortized compression, which struggles with out-of-distribution generalization, and Test-Time Training, which incurs prohibitive synthetic data costs and requires modifying model weights, creating stateful parameters that complicate concurrent serving. We propose Latent Context Compilation, a framework that fundamentally shifts context processing from adaptation to compilation. By utilizing a disposable LoRA module as a compiler, we distill long contexts into compact buffer tokens -- stateless, portable memory artifacts that are plug-and-play compatible with frozen base models. Crucially, we introduce a self-aligned optimization strategy that eliminates the need for synthetic context-relevant QA pairs. By regularizing context reconstruction task with context-agnostic random queries, we force compressed tokens to reside within the model's existing instruction-following manifold. Experiments with Llama-3.1-8B demonstrate that Latent Context Compilation preserves fine-grained details and reasoning capabilities where prior methods falter, effectively decoupling memory density from model parameters even at a 16x compression ratio.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21221
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Latent Context Compilation: Distilling Long Context into Compact Portable Memory
Li, Zeju
Zhou, Yizhou
Xu, Qiang
Machine Learning
Artificial Intelligence
Computation and Language
Efficient long-context LLM deployment is stalled by a dichotomy between amortized compression, which struggles with out-of-distribution generalization, and Test-Time Training, which incurs prohibitive synthetic data costs and requires modifying model weights, creating stateful parameters that complicate concurrent serving. We propose Latent Context Compilation, a framework that fundamentally shifts context processing from adaptation to compilation. By utilizing a disposable LoRA module as a compiler, we distill long contexts into compact buffer tokens -- stateless, portable memory artifacts that are plug-and-play compatible with frozen base models. Crucially, we introduce a self-aligned optimization strategy that eliminates the need for synthetic context-relevant QA pairs. By regularizing context reconstruction task with context-agnostic random queries, we force compressed tokens to reside within the model's existing instruction-following manifold. Experiments with Llama-3.1-8B demonstrate that Latent Context Compilation preserves fine-grained details and reasoning capabilities where prior methods falter, effectively decoupling memory density from model parameters even at a 16x compression ratio.
title Latent Context Compilation: Distilling Long Context into Compact Portable Memory
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.21221