Large Language Model as Token Compressor and Decompressor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenbing, Wang, Yiran, Song, Zikai, Zhang, Jielei, Zhao, Tianhao, Lin, Junkai, Yang, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916008366702592
author Li, Wenbing
Wang, Yiran
Song, Zikai
Zhang, Jielei
Zhao, Tianhao
Lin, Junkai
Yang, Wei
author_facet Li, Wenbing
Wang, Yiran
Song, Zikai
Zhang, Jielei
Zhao, Tianhao
Lin, Junkai
Yang, Wei
contents In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we design a self-expressive autoencoding framework that fine-tunes a pretrained LLM with lightweight LoRA adapters to map long texts into compact sequences of learned latent codes, termed Z-tokens, and to decode them back into natural language or task outputs. The resulting representation is content-adaptive: less predictable or information-dense segments can receive more Z-tokens, while redundant regions can be represented more compactly through a budget-aware length regularizer. Our method is evaluated on long-context datasets such as Wikipedia, CNN/DailyMail, HotpotQA, and QuALITY, showing that it preserves reconstruction quality and downstream performance while reducing effective context length, generation-stage memory usage, and end-to-end latency. This simple design supports both direct decoding from compressed contexts and autoregressive generation in the Z-token space, providing a practical interface for efficient long-context inference.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25340
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large Language Model as Token Compressor and Decompressor
Li, Wenbing
Wang, Yiran
Song, Zikai
Zhang, Jielei
Zhao, Tianhao
Lin, Junkai
Yang, Wei
Computation and Language
In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we design a self-expressive autoencoding framework that fine-tunes a pretrained LLM with lightweight LoRA adapters to map long texts into compact sequences of learned latent codes, termed Z-tokens, and to decode them back into natural language or task outputs. The resulting representation is content-adaptive: less predictable or information-dense segments can receive more Z-tokens, while redundant regions can be represented more compactly through a budget-aware length regularizer. Our method is evaluated on long-context datasets such as Wikipedia, CNN/DailyMail, HotpotQA, and QuALITY, showing that it preserves reconstruction quality and downstream performance while reducing effective context length, generation-stage memory usage, and end-to-end latency. This simple design supports both direct decoding from compressed contexts and autoregressive generation in the Z-token space, providing a practical interface for efficient long-context inference.
title Large Language Model as Token Compressor and Decompressor
topic Computation and Language
url https://arxiv.org/abs/2603.25340