Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ling, Zehui, Chen, Deshu, Zhang, Hongwei, Jiao, Yifeng, Guo, Xin, Cheng, Yuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909646948663296
author Ling, Zehui
Chen, Deshu
Zhang, Hongwei
Jiao, Yifeng
Guo, Xin
Cheng, Yuan
author_facet Ling, Zehui
Chen, Deshu
Zhang, Hongwei
Jiao, Yifeng
Guo, Xin
Cheng, Yuan
contents Large language models (LLMs) have demonstrated significant advancements in reasoning capabilities, performing well on various challenging benchmarks. Techniques like Chain-of-Thought prompting have been introduced to further improve reasoning. However, these approaches frequently generate longer outputs, which in turn increase computational latency. Although some methods use reinforcement learning to shorten reasoning, they often apply uniform penalties without considering the problem's complexity, leading to suboptimal outcomes. In this study, we seek to enhance the efficiency of LLM reasoning by promoting conciseness for simpler problems while preserving sufficient reasoning for more complex ones for accuracy, thus improving the model's overall performance. Specifically, we manage the model's reasoning efficiency by dividing the reward function and including a novel penalty for output length. Our approach has yielded impressive outcomes in benchmark evaluations across three datasets: GSM8K, MATH500, and AIME2024. For the comparatively simpler datasets GSM8K and MATH500, our method has effectively shortened output lengths while preserving or enhancing accuracy. On the more demanding AIME2024 dataset, our approach has resulted in improved accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
Ling, Zehui
Chen, Deshu
Zhang, Hongwei
Jiao, Yifeng
Guo, Xin
Cheng, Yuan
Computation and Language
Large language models (LLMs) have demonstrated significant advancements in reasoning capabilities, performing well on various challenging benchmarks. Techniques like Chain-of-Thought prompting have been introduced to further improve reasoning. However, these approaches frequently generate longer outputs, which in turn increase computational latency. Although some methods use reinforcement learning to shorten reasoning, they often apply uniform penalties without considering the problem's complexity, leading to suboptimal outcomes. In this study, we seek to enhance the efficiency of LLM reasoning by promoting conciseness for simpler problems while preserving sufficient reasoning for more complex ones for accuracy, thus improving the model's overall performance. Specifically, we manage the model's reasoning efficiency by dividing the reward function and including a novel penalty for output length. Our approach has yielded impressive outcomes in benchmark evaluations across three datasets: GSM8K, MATH500, and AIME2024. For the comparatively simpler datasets GSM8K and MATH500, our method has effectively shortened output lengths while preserving or enhancing accuracy. On the more demanding AIME2024 dataset, our approach has resulted in improved accuracy.
title Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
topic Computation and Language
url https://arxiv.org/abs/2506.10446