LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xingyu, Yan, Yuchen, Lyu, Shangke, Wu, Linjuan, Qiu, Yiwen, Shen, Yongliang, Lu, Weiming, Shao, Jian, Xiao, Jun, Zhuang, Yueting
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908489291399168
author Wu, Xingyu
Yan, Yuchen
Lyu, Shangke
Wu, Linjuan
Qiu, Yiwen
Shen, Yongliang
Lu, Weiming
Shao, Jian
Xiao, Jun
Zhuang, Yueting
author_facet Wu, Xingyu
Yan, Yuchen
Lyu, Shangke
Wu, Linjuan
Qiu, Yiwen
Shen, Yongliang
Lu, Weiming
Shao, Jian
Xiao, Jun
Zhuang, Yueting
contents Large reasoning models have achieved remarkable performance through extended chain-of-thought sequences, yet this computational freedom leads to excessive token generation even for simple problems. We present Length-Adaptive Policy Optimization (LAPO), a novel framework that transforms reasoning length control from an external constraint into an intrinsic model capability. Unlike existing approaches that impose rigid limits or rely on post-hoc interventions, LAPO enables models to internalize an understanding of appropriate reasoning depth through a two-stage reinforcement learning process. In the first stage, models learn natural reasoning patterns by discovering the statistical distribution of successful solution lengths. The second stage leverages these patterns as meta-cognitive guidance, embedding them directly within the model's reasoning context to ensure inference-time flexibility. Experiments on mathematical reasoning benchmarks demonstrate that LAPO reduces token usage by up to 40.9% while improving accuracy by 2.3%. Our analysis reveals that models trained with LAPO develop emergent abilities to allocate computational resources based on problem complexity, achieving efficient reasoning without sacrificing quality.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
Wu, Xingyu
Yan, Yuchen
Lyu, Shangke
Wu, Linjuan
Qiu, Yiwen
Shen, Yongliang
Lu, Weiming
Shao, Jian
Xiao, Jun
Zhuang, Yueting
Artificial Intelligence
Computation and Language
Large reasoning models have achieved remarkable performance through extended chain-of-thought sequences, yet this computational freedom leads to excessive token generation even for simple problems. We present Length-Adaptive Policy Optimization (LAPO), a novel framework that transforms reasoning length control from an external constraint into an intrinsic model capability. Unlike existing approaches that impose rigid limits or rely on post-hoc interventions, LAPO enables models to internalize an understanding of appropriate reasoning depth through a two-stage reinforcement learning process. In the first stage, models learn natural reasoning patterns by discovering the statistical distribution of successful solution lengths. The second stage leverages these patterns as meta-cognitive guidance, embedding them directly within the model's reasoning context to ensure inference-time flexibility. Experiments on mathematical reasoning benchmarks demonstrate that LAPO reduces token usage by up to 40.9% while improving accuracy by 2.3%. Our analysis reveals that models trained with LAPO develop emergent abilities to allocate computational resources based on problem complexity, achieving efficient reasoning without sacrificing quality.
title LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.15758