KL-Regularized Reinforcement Learning is Designed to Mode Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: GX-Chen, Anthony, Prakash, Jatin, Guo, Jeff, Fergus, Rob, Ranganath, Rajesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911228472852480
author GX-Chen, Anthony
Prakash, Jatin
Guo, Jeff
Fergus, Rob
Ranganath, Rajesh
author_facet GX-Chen, Anthony
Prakash, Jatin
Guo, Jeff
Fergus, Rob
Ranganath, Rajesh
contents It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred if the goal is to sample from multiple diverse modes. We show -- mathematically and empirically -- that this intuition does not necessarily transfer well to doing reinforcement learning with reverse/forward KL regularization (e.g. as commonly used with language models). Instead, the choice of reverse/forward KL determines the family of optimal target distributions, parameterized by the regularization coefficient. Mode coverage depends primarily on other factors, such as regularization strength, and relative scales between rewards and reference probabilities. Further, we show commonly used settings such as low regularization strength and equal verifiable rewards tend to specify unimodal target distributions, meaning the optimization objective is, by construction, non-diverse. We leverage these insights to construct a simple, scalable, and theoretically justified algorithm. It makes minimal changes to reward magnitudes, yet optimizes for a target distribution which puts high probability over all high-quality sampling modes. In experiments, this simple modification works to post-train both Large Language Models and Chemical Language Models to have higher solution quality and diversity, without any external signals of diversity, and works with both forward and reverse KL when using either naively fails.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KL-Regularized Reinforcement Learning is Designed to Mode Collapse
GX-Chen, Anthony
Prakash, Jatin
Guo, Jeff
Fergus, Rob
Ranganath, Rajesh
Machine Learning
It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred if the goal is to sample from multiple diverse modes. We show -- mathematically and empirically -- that this intuition does not necessarily transfer well to doing reinforcement learning with reverse/forward KL regularization (e.g. as commonly used with language models). Instead, the choice of reverse/forward KL determines the family of optimal target distributions, parameterized by the regularization coefficient. Mode coverage depends primarily on other factors, such as regularization strength, and relative scales between rewards and reference probabilities. Further, we show commonly used settings such as low regularization strength and equal verifiable rewards tend to specify unimodal target distributions, meaning the optimization objective is, by construction, non-diverse. We leverage these insights to construct a simple, scalable, and theoretically justified algorithm. It makes minimal changes to reward magnitudes, yet optimizes for a target distribution which puts high probability over all high-quality sampling modes. In experiments, this simple modification works to post-train both Large Language Models and Chemical Language Models to have higher solution quality and diversity, without any external signals of diversity, and works with both forward and reverse KL when using either naively fails.
title KL-Regularized Reinforcement Learning is Designed to Mode Collapse
topic Machine Learning
url https://arxiv.org/abs/2510.20817