DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Tian, Jiao, Wenxiang, He, Zhiwei, Xu, Jiahao, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908906294345728
author Liang, Tian
Jiao, Wenxiang
He, Zhiwei
Xu, Jiahao
Mi, Haitao
Yu, Dong
author_facet Liang, Tian
Jiao, Wenxiang
He, Zhiwei
Xu, Jiahao
Mi, Haitao
Yu, Dong
contents Large Reasoning Models (LRMs) have demonstrated impressive capabilities but suffer from cognitive inefficiencies like "overthinking" simple problems and "underthinking" complex ones. While existing methods that use supervised fine-tuning (SFT) or reinforcement learning (RL) with token-length rewards can improve efficiency, they often do so at the cost of accuracy. This paper introduces DeepCompress, a novel framework that simultaneously enhances both the accuracy and efficiency of LRMs. We challenge the prevailing approach of consistently favoring shorter reasoning paths, showing that longer responses can contain a broader range of correct solutions for difficult problems. DeepCompress employs an adaptive length reward mechanism that dynamically classifies problems as "Simple" or "Hard" in real-time based on the model's evolving capability. It encourages shorter, more efficient reasoning for "Simple" problems while promoting longer, more exploratory thought chains for "Hard" problems. This dual-reward strategy enables the model to autonomously adjust its Chain-of-Thought (CoT) length, compressing reasoning for well-mastered problems and extending it for those it finds challenging. Experimental results on challenging mathematical benchmarks show that DeepCompress consistently outperforms baseline methods, achieving superior accuracy while significantly improving token efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
Liang, Tian
Jiao, Wenxiang
He, Zhiwei
Xu, Jiahao
Mi, Haitao
Yu, Dong
Artificial Intelligence
Computation and Language
Large Reasoning Models (LRMs) have demonstrated impressive capabilities but suffer from cognitive inefficiencies like "overthinking" simple problems and "underthinking" complex ones. While existing methods that use supervised fine-tuning (SFT) or reinforcement learning (RL) with token-length rewards can improve efficiency, they often do so at the cost of accuracy. This paper introduces DeepCompress, a novel framework that simultaneously enhances both the accuracy and efficiency of LRMs. We challenge the prevailing approach of consistently favoring shorter reasoning paths, showing that longer responses can contain a broader range of correct solutions for difficult problems. DeepCompress employs an adaptive length reward mechanism that dynamically classifies problems as "Simple" or "Hard" in real-time based on the model's evolving capability. It encourages shorter, more efficient reasoning for "Simple" problems while promoting longer, more exploratory thought chains for "Hard" problems. This dual-reward strategy enables the model to autonomously adjust its Chain-of-Thought (CoT) length, compressing reasoning for well-mastered problems and extending it for those it finds challenging. Experimental results on challenging mathematical benchmarks show that DeepCompress consistently outperforms baseline methods, achieving superior accuracy while significantly improving token efficiency.
title DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.27419