Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yiqun, Feng, Jinyuan, Yang, Wei, Zhong, Meizhi, Shi, Zhengliang, Li, Rui, Wei, Xiaochi, Gao, Yan, Wu, Yi, Hu, Yao, Pu, Zhiqiang, Mao, Jiaxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908798251171840
author Chen, Yiqun
Feng, Jinyuan
Yang, Wei
Zhong, Meizhi
Shi, Zhengliang
Li, Rui
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
Mao, Jiaxin
author_facet Chen, Yiqun
Feng, Jinyuan
Yang, Wei
Zhong, Meizhi
Shi, Zhengliang
Li, Rui
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
Mao, Jiaxin
contents The inference overhead induced by redundant reasoning undermines the interactive experience and severely bottlenecks the deployment of Large Reasoning Models. Existing reinforcement learning (RL)-based solutions tackle this problem by coupling a length penalty with outcome-based rewards. This simplistic reward weighting struggles to reconcile brevity with accuracy, as enforcing brevity may compromise critical reasoning logic. In this work, we address this limitation by proposing a multi-agent RL framework that selectively penalizes redundant chunks, while preserving essential reasoning logic. Our framework, Self-Compression via MARL (SCMA), instantiates redundancy detection and evaluation through two specialized agents: \textbf{a Segmentation Agent} for decomposing the reasoning process into logical chunks, and \textbf{a Scoring Agent} for quantifying the significance of each chunk. The Segmentation and Scoring agents collaboratively define an importance-weighted length penalty during training, incentivizing \textbf{a Reasoning Agent} to prioritize essential logic without introducing inference overhead during deployment. Empirical evaluations across model scales demonstrate that SCMA reduces response length by 11.1\% to 39.0\% while boosting accuracy by 4.33\% to 10.02\%. Furthermore, ablation studies and qualitative analysis validate that the synergistic optimization within the MARL framework fosters emergent behaviors, yielding more powerful LRMs compared to vanilla RL paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21919
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
Chen, Yiqun
Feng, Jinyuan
Yang, Wei
Zhong, Meizhi
Shi, Zhengliang
Li, Rui
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
Mao, Jiaxin
Artificial Intelligence
Computation and Language
The inference overhead induced by redundant reasoning undermines the interactive experience and severely bottlenecks the deployment of Large Reasoning Models. Existing reinforcement learning (RL)-based solutions tackle this problem by coupling a length penalty with outcome-based rewards. This simplistic reward weighting struggles to reconcile brevity with accuracy, as enforcing brevity may compromise critical reasoning logic. In this work, we address this limitation by proposing a multi-agent RL framework that selectively penalizes redundant chunks, while preserving essential reasoning logic. Our framework, Self-Compression via MARL (SCMA), instantiates redundancy detection and evaluation through two specialized agents: \textbf{a Segmentation Agent} for decomposing the reasoning process into logical chunks, and \textbf{a Scoring Agent} for quantifying the significance of each chunk. The Segmentation and Scoring agents collaboratively define an importance-weighted length penalty during training, incentivizing \textbf{a Reasoning Agent} to prioritize essential logic without introducing inference overhead during deployment. Empirical evaluations across model scales demonstrate that SCMA reduces response length by 11.1\% to 39.0\% while boosting accuracy by 4.33\% to 10.02\%. Furthermore, ablation studies and qualitative analysis validate that the synergistic optimization within the MARL framework fosters emergent behaviors, yielding more powerful LRMs compared to vanilla RL paradigms.
title Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.21919