FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Chiyu, Yang, Shuo, Huang, Kexin, Lu, Jinda, Meng, Haoming, Wang, Shangshang, Ding, Bolin, Vosoughi, Soroush, Wang, Guoyin, Zhou, Jingren
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911555320283136
author Ma, Chiyu
Yang, Shuo
Huang, Kexin
Lu, Jinda
Meng, Haoming
Wang, Shangshang
Ding, Bolin
Vosoughi, Soroush
Wang, Guoyin
Zhou, Jingren
author_facet Ma, Chiyu
Yang, Shuo
Huang, Kexin
Lu, Jinda
Meng, Haoming
Wang, Shangshang
Ding, Bolin
Vosoughi, Soroush
Wang, Guoyin
Zhou, Jingren
contents We present Future-KL Influenced Policy Optimization (FIPO), a reinforcement learning algorithm designed to overcome reasoning bottlenecks in large language models. While GRPO style training scales effectively, it typically relies on outcome-based rewards (ORM) that distribute a global advantage uniformly across every token in a trajectory. We argue that this coarse-grained credit assignment imposes a performance ceiling by failing to distinguish critical logical pivots from trivial tokens. FIPO addresses this by incorporating discounted future-KL divergence into the policy update, creating a dense advantage formulation that re-weights tokens based on their influence on subsequent trajectory behavior. Empirically, FIPO enables models to break through the length stagnation seen in standard baselines. Evaluated on Qwen2.5-32B, FIPO extends the average chain-of-thought length from roughly 4,000 to over 10,000 tokens and increases AIME 2024 Pass@1 accuracy from 50.0% to a peak of 58.0% (converging at approximately 56.0\%). This outperforms both DeepSeek-R1-Zero-Math-32B (around 47.0%) and o1-mini (approximately 56.0%). Our results suggest that establishing dense advantage formulations is a vital path for evolving ORM-based algorithms to unlock the full reasoning potential of base models. We open-source our training system, built on the verl framework.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19835
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
Ma, Chiyu
Yang, Shuo
Huang, Kexin
Lu, Jinda
Meng, Haoming
Wang, Shangshang
Ding, Bolin
Vosoughi, Soroush
Wang, Guoyin
Zhou, Jingren
Machine Learning
We present Future-KL Influenced Policy Optimization (FIPO), a reinforcement learning algorithm designed to overcome reasoning bottlenecks in large language models. While GRPO style training scales effectively, it typically relies on outcome-based rewards (ORM) that distribute a global advantage uniformly across every token in a trajectory. We argue that this coarse-grained credit assignment imposes a performance ceiling by failing to distinguish critical logical pivots from trivial tokens. FIPO addresses this by incorporating discounted future-KL divergence into the policy update, creating a dense advantage formulation that re-weights tokens based on their influence on subsequent trajectory behavior. Empirically, FIPO enables models to break through the length stagnation seen in standard baselines. Evaluated on Qwen2.5-32B, FIPO extends the average chain-of-thought length from roughly 4,000 to over 10,000 tokens and increases AIME 2024 Pass@1 accuracy from 50.0% to a peak of 58.0% (converging at approximately 56.0\%). This outperforms both DeepSeek-R1-Zero-Math-32B (around 47.0%) and o1-mini (approximately 56.0%). Our results suggest that establishing dense advantage formulations is a vital path for evolving ORM-based algorithms to unlock the full reasoning potential of base models. We open-source our training system, built on the verl framework.
title FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
topic Machine Learning
url https://arxiv.org/abs/2603.19835