ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Deng, Jie, Liang, Shining, Li, Jun, Li, Hongzhi, Xie, Yutao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910008358207488
author Deng, Jie
Liang, Shining
Li, Jun
Li, Hongzhi
Xie, Yutao
author_facet Deng, Jie
Liang, Shining
Li, Jun
Li, Hongzhi
Xie, Yutao
contents Large reasoning models (LRMs) typically solve reasoning-intensive tasks by generating long chain-of-thought (CoT) traces, leading to substantial inference overhead. We identify a reproducible inference-time phenomenon, termed Self-Compression: when multiple independent and answerable questions are presented within a single prompt, the model spontaneously produces shorter reasoning traces for each question. This phenomenon arises from multi-question contextual pressure during generation and consistently manifests across models and benchmarks. Building on this observation, we propose ConPress (Learning from Contextual Pressure), a lightweight self-supervised fine-tuning approach. ConPress constructs multi-question prompts to induce self-compression, samples the resulting model outputs, and parses and filters per-question traces to obtain concise yet correct reasoning trajectories. These trajectories are directly used for supervised fine-tuning, internalizing compressed reasoning behavior in single-question settings without external teachers, manual pruning, or reinforcement learning. With only 8k fine-tuning examples, ConPress reduces reasoning token usage by 59% on MATH500 and 33% on AIME25, while maintaining competitive accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01472
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
Deng, Jie
Liang, Shining
Li, Jun
Li, Hongzhi
Xie, Yutao
Computation and Language
Large reasoning models (LRMs) typically solve reasoning-intensive tasks by generating long chain-of-thought (CoT) traces, leading to substantial inference overhead. We identify a reproducible inference-time phenomenon, termed Self-Compression: when multiple independent and answerable questions are presented within a single prompt, the model spontaneously produces shorter reasoning traces for each question. This phenomenon arises from multi-question contextual pressure during generation and consistently manifests across models and benchmarks. Building on this observation, we propose ConPress (Learning from Contextual Pressure), a lightweight self-supervised fine-tuning approach. ConPress constructs multi-question prompts to induce self-compression, samples the resulting model outputs, and parses and filters per-question traces to obtain concise yet correct reasoning trajectories. These trajectories are directly used for supervised fine-tuning, internalizing compressed reasoning behavior in single-question settings without external teachers, manual pruning, or reinforcement learning. With only 8k fine-tuning examples, ConPress reduces reasoning token usage by 59% on MATH500 and 33% on AIME25, while maintaining competitive accuracy.
title ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
topic Computation and Language
url https://arxiv.org/abs/2602.01472