Variational Reasoning for Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Xiangxin, Liu, Zichen, Wang, Haonan, Du, Chao, Lin, Min, Li, Chongxuan, Wang, Liang, Pang, Tianyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908595134660608
author Zhou, Xiangxin
Liu, Zichen
Wang, Haonan
Du, Chao
Lin, Min
Li, Chongxuan
Wang, Liang
Pang, Tianyu
author_facet Zhou, Xiangxin
Liu, Zichen
Wang, Haonan
Du, Chao
Lin, Min
Li, Chongxuan
Wang, Liang
Pang, Tianyu
contents We introduce a variational reasoning framework for language models that treats thinking traces as latent variables and optimizes them through variational inference. Starting from the evidence lower bound (ELBO), we extend it to a multi-trace objective for tighter bounds and propose a forward-KL formulation that stabilizes the training of the variational posterior. We further show that rejection sampling finetuning and binary-reward RL, including GRPO, can be interpreted as local forward-KL objectives, where an implicit weighting by model accuracy naturally arises from the derivation and reveals a previously unnoticed bias toward easier questions. We empirically validate our method on the Qwen 2.5 and Qwen 3 model families across a wide range of reasoning tasks. Overall, our work provides a principled probabilistic perspective that unifies variational inference with RL-style methods and yields stable objectives for improving the reasoning ability of language models. Our code is available at https://github.com/sail-sg/variational-reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22637
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Variational Reasoning for Language Models
Zhou, Xiangxin
Liu, Zichen
Wang, Haonan
Du, Chao
Lin, Min
Li, Chongxuan
Wang, Liang
Pang, Tianyu
Computation and Language
Artificial Intelligence
Machine Learning
We introduce a variational reasoning framework for language models that treats thinking traces as latent variables and optimizes them through variational inference. Starting from the evidence lower bound (ELBO), we extend it to a multi-trace objective for tighter bounds and propose a forward-KL formulation that stabilizes the training of the variational posterior. We further show that rejection sampling finetuning and binary-reward RL, including GRPO, can be interpreted as local forward-KL objectives, where an implicit weighting by model accuracy naturally arises from the derivation and reveals a previously unnoticed bias toward easier questions. We empirically validate our method on the Qwen 2.5 and Qwen 3 model families across a wide range of reasoning tasks. Overall, our work provides a principled probabilistic perspective that unifies variational inference with RL-style methods and yields stable objectives for improving the reasoning ability of language models. Our code is available at https://github.com/sail-sg/variational-reasoning.
title Variational Reasoning for Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.22637