Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Guoxin, Liu, Yibing, Li, Chengzhengxu, Liang, Yu, Wang, Yan, Zhang, Yueyang, Chen, Kecheng, Zhang, Zhaohan, Sun, Zhiyuan, Shi, Daiting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911725142409216
author Ma, Guoxin
Liu, Yibing
Li, Chengzhengxu
Liang, Yu
Wang, Yan
Zhang, Yueyang
Chen, Kecheng
Zhang, Zhaohan
Sun, Zhiyuan
Shi, Daiting
author_facet Ma, Guoxin
Liu, Yibing
Li, Chengzhengxu
Liang, Yu
Wang, Yan
Zhang, Yueyang
Chen, Kecheng
Zhang, Zhaohan
Sun, Zhiyuan
Shi, Daiting
contents Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-specific training, leaving the intrinsic capabilities of LLMs underexplored. In contrast, this work reveals that a thinking model itself can naturally compress long contexts by organizing task-relevant information. We thus derive Thinking as Compression (TaC), a new compression paradigm that treats thinking itself as compressed context. Without relying on specific dedicated compressor, TaC directly prompts the thinking model to generate thinking traces as the shortened context, already outperforming most representative compression methods. Further, given that raw thinking output may struggle with budget control and shortcut behaviors, we introduce Thinking as Compression Constrained (TaC-C), leveraging a simple reward-driven optimization framework to elicit intrinsic thinking as compact and controllable compressed context. Experiments across four long-context QA benchmarks demonstrate that TaC-C consistently outperforms existing baselines. At 4x and 8x compression ratios, it surpasses the strongest competitor by 17.4% and 23.4% in average F1, and by 15.7% and 21.7% in average Exact Match Score (EM), respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28713
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
Ma, Guoxin
Liu, Yibing
Li, Chengzhengxu
Liang, Yu
Wang, Yan
Zhang, Yueyang
Chen, Kecheng
Zhang, Zhaohan
Sun, Zhiyuan
Shi, Daiting
Artificial Intelligence
Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-specific training, leaving the intrinsic capabilities of LLMs underexplored. In contrast, this work reveals that a thinking model itself can naturally compress long contexts by organizing task-relevant information. We thus derive Thinking as Compression (TaC), a new compression paradigm that treats thinking itself as compressed context. Without relying on specific dedicated compressor, TaC directly prompts the thinking model to generate thinking traces as the shortened context, already outperforming most representative compression methods. Further, given that raw thinking output may struggle with budget control and shortcut behaviors, we introduce Thinking as Compression Constrained (TaC-C), leveraging a simple reward-driven optimization framework to elicit intrinsic thinking as compact and controllable compressed context. Experiments across four long-context QA benchmarks demonstrate that TaC-C consistently outperforms existing baselines. At 4x and 8x compression ratios, it surpasses the strongest competitor by 17.4% and 23.4% in average F1, and by 15.7% and 21.7% in average Exact Match Score (EM), respectively.
title Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
topic Artificial Intelligence
url https://arxiv.org/abs/2605.28713