Coupled Variational Reinforcement Learning for Language Model General Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wen, Xueru, Lou, Jie, Liu, Yanjiang, Lin, Hongyu, He, Ben, Han, Xianpei, Sun, Le, Lu, Yaojie, Zhang, Debing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913157567479808
author Wen, Xueru
Lou, Jie
Liu, Yanjiang
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Lu, Yaojie
Zhang, Debing
author_facet Wen, Xueru
Lou, Jie
Liu, Yanjiang
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Lu, Yaojie
Zhang, Debing
contents While reinforcement learning has achieved impressive progress in language model reasoning, it is constrained by the requirement for verifiable rewards. Recent verifier-free RL methods address this limitation by utilizing the probabilities that LLMs generate reference answers as reward signals. However, these approaches typically sample reasoning traces conditioned only on the question. This design decouples reasoning-trace sampling from answer information, leading to inefficient exploration and incoherence between traces and final answers. In this paper, we propose \textit{\b{Co}upled \b{V}ariational \b{R}einforcement \b{L}earning} (CoVRL), which bridges variational inference and reinforcement learning by coupling prior and posterior distributions through a hybrid sampling strategy. By constructing and optimizing a composite distribution that integrates these two distributions, CoVRL enables efficient exploration while preserving strong thought-answer coherence. Extensive experiments on mathematical and general reasoning benchmarks show that CoVRL improves performance by 12.4\% over the base model and achieves an additional 2.3\% improvement over state-of-the-art verifier-free RL baselines, providing a principled framework for enhancing the general reasoning capabilities of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Coupled Variational Reinforcement Learning for Language Model General Reasoning
Wen, Xueru
Lou, Jie
Liu, Yanjiang
Lin, Hongyu
He, Ben
Han, Xianpei
Sun, Le
Lu, Yaojie
Zhang, Debing
Computation and Language
Artificial Intelligence
While reinforcement learning has achieved impressive progress in language model reasoning, it is constrained by the requirement for verifiable rewards. Recent verifier-free RL methods address this limitation by utilizing the probabilities that LLMs generate reference answers as reward signals. However, these approaches typically sample reasoning traces conditioned only on the question. This design decouples reasoning-trace sampling from answer information, leading to inefficient exploration and incoherence between traces and final answers. In this paper, we propose \textit{\b{Co}upled \b{V}ariational \b{R}einforcement \b{L}earning} (CoVRL), which bridges variational inference and reinforcement learning by coupling prior and posterior distributions through a hybrid sampling strategy. By constructing and optimizing a composite distribution that integrates these two distributions, CoVRL enables efficient exploration while preserving strong thought-answer coherence. Extensive experiments on mathematical and general reasoning benchmarks show that CoVRL improves performance by 12.4\% over the base model and achieves an additional 2.3\% improvement over state-of-the-art verifier-free RL baselines, providing a principled framework for enhancing the general reasoning capabilities of language models.
title Coupled Variational Reinforcement Learning for Language Model General Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.12576