Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Xinchen, Afifi, Hossam, Marot, Michel, Wang, Xilu, Yin, Lu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914383881306112
author Han, Xinchen
Afifi, Hossam
Marot, Michel
Wang, Xilu
Yin, Lu
author_facet Han, Xinchen
Afifi, Hossam
Marot, Michel
Wang, Xilu
Yin, Lu
contents Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency without proportional performance gains. In this paper, we propose Fine-grained Group policy Optimization (FGO), a Reinforcement Learning (RL) algorithm that refines group responses by subdividing them and assigning appropriate weights based on length and entropy, thereby enabling effective CoT compression. Meanwhile, as an enhanced variant of Group Relative Policy Optimization (GRPO), FGO successfully addresses two major limitations of the GRPO: inefficient data utilization and entropy collapse. We evaluate FGO on multiple reasoning LLMs and benchmarks, including MATH500, AIME24, AMC23, and Minerva. Experimental results show that FGO achieves efficient CoT compression without degrading performance, and simultaneously resolves the key limitations of GRPO. Code: https://github.com/Mr-XcHan/FGO.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10048
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
Han, Xinchen
Afifi, Hossam
Marot, Michel
Wang, Xilu
Yin, Lu
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency without proportional performance gains. In this paper, we propose Fine-grained Group policy Optimization (FGO), a Reinforcement Learning (RL) algorithm that refines group responses by subdividing them and assigning appropriate weights based on length and entropy, thereby enabling effective CoT compression. Meanwhile, as an enhanced variant of Group Relative Policy Optimization (GRPO), FGO successfully addresses two major limitations of the GRPO: inefficient data utilization and entropy collapse. We evaluate FGO on multiple reasoning LLMs and benchmarks, including MATH500, AIME24, AMC23, and Minerva. Experimental results show that FGO achieves efficient CoT compression without degrading performance, and simultaneously resolves the key limitations of GRPO. Code: https://github.com/Mr-XcHan/FGO.
title Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.10048