ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dumitru, Razvan-Gabriel, Peteleaza, Darius, Yadav, Vikas, Pan, Liangming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908376151097344
author Dumitru, Razvan-Gabriel
Peteleaza, Darius
Yadav, Vikas
Pan, Liangming
author_facet Dumitru, Razvan-Gabriel
Peteleaza, Darius
Yadav, Vikas
Pan, Liangming
contents Large language models excel at complex tasks by breaking down problems into structured reasoning steps. However, reasoning traces often extend beyond reaching a correct answer, causing wasted computation, reduced readability, and hallucinations. To address this, we introduce a novel hyperparameter-free conciseness score used as a reward signal within a reinforcement learning framework to guide models toward generating correct and concise reasoning traces. This score is evaluated by a large language model acting as a judge, enabling dynamic, context-aware feedback beyond simple token length. Our method achieves state-of-the-art efficiency-accuracy trade-offs on the MATH dataset, reducing token usage by up to 31x on simple problems while improving accuracy by 7%, and on the hardest problems, it outperforms full reasoning by +7.5% accuracy with up to 3.6x fewer tokens. On TheoremQA, our method improves accuracy by +2.2% using 12.5x fewer tokens. We also conduct ablation studies on the judge model, reward composition, and problem difficulty, showing that our method dynamically adapts reasoning length based on problem difficulty and benefits significantly from stronger judges. The code, model weights, and datasets are open-sourced at https://github.com/RazvanDu/ConciseRL.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17250
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
Dumitru, Razvan-Gabriel
Peteleaza, Darius
Yadav, Vikas
Pan, Liangming
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.0
Large language models excel at complex tasks by breaking down problems into structured reasoning steps. However, reasoning traces often extend beyond reaching a correct answer, causing wasted computation, reduced readability, and hallucinations. To address this, we introduce a novel hyperparameter-free conciseness score used as a reward signal within a reinforcement learning framework to guide models toward generating correct and concise reasoning traces. This score is evaluated by a large language model acting as a judge, enabling dynamic, context-aware feedback beyond simple token length. Our method achieves state-of-the-art efficiency-accuracy trade-offs on the MATH dataset, reducing token usage by up to 31x on simple problems while improving accuracy by 7%, and on the hardest problems, it outperforms full reasoning by +7.5% accuracy with up to 3.6x fewer tokens. On TheoremQA, our method improves accuracy by +2.2% using 12.5x fewer tokens. We also conduct ablation studies on the judge model, reward composition, and problem difficulty, showing that our method dynamically adapts reasoning length based on problem difficulty and benefits significantly from stronger judges. The code, model weights, and datasets are open-sourced at https://github.com/RazvanDu/ConciseRL.
title ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.0
url https://arxiv.org/abs/2505.17250