Towards Flash Thinking via Decoupled Advantage Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tan, Zezhong, Gao, Hang, Ma, Xinhong, Zhang, Feng, Dong, Ziqiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911216451977216
author Tan, Zezhong
Gao, Hang
Ma, Xinhong
Zhang, Feng
Dong, Ziqiang
author_facet Tan, Zezhong
Gao, Hang
Ma, Xinhong
Zhang, Feng
Dong, Ziqiang
contents Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although existing RL algorithms significantly enhance model accuracy, they still suffer from excessively lengthy responses and overthinking issues, resulting in increased inference latency and computational consumption, especially for simple tasks that require minimal reasoning. To address this, we propose a novel RL framework, DEPO, to reduce inefficient reasoning for models. Our method mainly consists of three core components: (1) an innovative advantage decoupled algorithm to guide model reduction of inefficient tokens; (2) a difficulty-aware length penalty to lower the overall length of model responses; (3) an advantage clipping method to prevent bias in policy optimization. In our experiments, applied to DeepSeek-Distill-Qwen-7B and DeepSeek-Distill-Qwen-1.5B as base models, DEPO achieves a significant reduction in sequence length by 39% and reduces excessive reasoning paths in inefficient tokens, while outperforming the base model in overall accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Flash Thinking via Decoupled Advantage Policy Optimization
Tan, Zezhong
Gao, Hang
Ma, Xinhong
Zhang, Feng
Dong, Ziqiang
Artificial Intelligence
Machine Learning
Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although existing RL algorithms significantly enhance model accuracy, they still suffer from excessively lengthy responses and overthinking issues, resulting in increased inference latency and computational consumption, especially for simple tasks that require minimal reasoning. To address this, we propose a novel RL framework, DEPO, to reduce inefficient reasoning for models. Our method mainly consists of three core components: (1) an innovative advantage decoupled algorithm to guide model reduction of inefficient tokens; (2) a difficulty-aware length penalty to lower the overall length of model responses; (3) an advantage clipping method to prevent bias in policy optimization. In our experiments, applied to DeepSeek-Distill-Qwen-7B and DeepSeek-Distill-Qwen-1.5B as base models, DEPO achieves a significant reduction in sequence length by 39% and reduces excessive reasoning paths in inefficient tokens, while outperforming the base model in overall accuracy.
title Towards Flash Thinking via Decoupled Advantage Policy Optimization
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.15374