PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Ruishuo, Chen, Yu, Li, Zhuoran, Huang, Longbo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916045110902784
author Chen, Ruishuo
Chen, Yu
Li, Zhuoran
Huang, Longbo
author_facet Chen, Ruishuo
Chen, Yu
Li, Zhuoran
Huang, Longbo
contents Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting $α$-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution ($α> 1$) to intensify logical reasoning, or flattening it ($α< 1$) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18363
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Chen, Ruishuo
Chen, Yu
Li, Zhuoran
Huang, Longbo
Computation and Language
Artificial Intelligence
Machine Learning
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting $α$-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution ($α> 1$) to intensify logical reasoning, or flattening it ($α< 1$) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.
title PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.18363