ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yiling, Wei, Hongchen, Chen, Zhenzhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914620286959616
author Gao, Yiling
Wei, Hongchen
Chen, Zhenzhong
author_facet Gao, Yiling
Wei, Hongchen
Chen, Zhenzhong
contents In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To address this problem, we propose an Extreme Token Compression (ETC) framework that minimizes task loss when reducing the number of input tokens based on the principle of variational information distillation. Specifically, from an information-theoretic perspective, we show that minimizing task loss requires the compact representation to preserve the instruction-aware sufficient statistic of the task-relevant visual information for prediction. In practice, ETC leverages text-to-image cross-attention to weight the original visual features to approximate the latent instruction-aware predictive statistic. Moreover, ETC introduces a variational information distillation, enabling the compact representation to preserve the essential information to recover this predictive statistic. Experiments on LLaVA-1.5-7B and Qwen3-VL-2B show that ETC remains effective even under single-token compression, substantially reducing KV-cache overhead while retaining strong task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
Gao, Yiling
Wei, Hongchen
Chen, Zhenzhong
Computer Vision and Pattern Recognition
In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To address this problem, we propose an Extreme Token Compression (ETC) framework that minimizes task loss when reducing the number of input tokens based on the principle of variational information distillation. Specifically, from an information-theoretic perspective, we show that minimizing task loss requires the compact representation to preserve the instruction-aware sufficient statistic of the task-relevant visual information for prediction. In practice, ETC leverages text-to-image cross-attention to weight the original visual features to approximate the latent instruction-aware predictive statistic. Moreover, ETC introduces a variational information distillation, enabling the compact representation to preserve the essential information to recover this predictive statistic. Experiments on LLaVA-1.5-7B and Qwen3-VL-2B show that ETC remains effective even under single-token compression, substantially reducing KV-cache overhead while retaining strong task performance.
title ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.00543