KL-Regularized Reinforcement Learning is Designed to Mode Collapse
Fuente:
arXiv
Saved in:
| Main Authors: | GX-Chen, Anthony, Prakash, Jatin, Guo, Jeff, Fergus, Rob, Ranganath, Rajesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction
by: GX-Chen, Anthony, et al.
Published: (2024)
by: GX-Chen, Anthony, et al.
Published: (2024)
Attention and Compression is all you need for Controllably Efficient Language Models
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026)
Light-weight probing of unsupervised representations for Reinforcement Learning
by: Zhang, Wancong, et al.
Published: (2022)
by: Zhang, Wancong, et al.
Published: (2022)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
by: Dragutinović, Sara, et al.
Published: (2026)
by: Dragutinović, Sara, et al.
Published: (2026)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Preference learning made easy: Everything should be understood through win rate
by: Zhang, Lily H., et al.
Published: (2025)
by: Zhang, Lily H., et al.
Published: (2025)
What's the score? Automated Denoising Score Matching for Nonlinear Diffusions
by: Singhal, Raghav, et al.
Published: (2024)
by: Singhal, Raghav, et al.
Published: (2024)
What Can You Do When You Have Zero Rewards During RL?
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024)
by: Puli, Aahlad, et al.
Published: (2024)
Stochastic Video Generation with a Learned Prior
by: Denton, Remi, et al.
Published: (2018)
by: Denton, Remi, et al.
Published: (2018)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
by: Piterbarg, Ulyana, et al.
Published: (2024)
by: Piterbarg, Ulyana, et al.
Published: (2024)
HingeRLC-GAN: Combating Mode Collapse with Hinge Loss and RLC Regularization
by: Goni, Osman, et al.
Published: (2025)
by: Goni, Osman, et al.
Published: (2025)
Symmetric Rank-One Quasi-Newton Methods for Deep Learning Using Cubic Regularization
by: Ranganath, Aditya, et al.
Published: (2025)
by: Ranganath, Aditya, et al.
Published: (2025)
Reinforcement Learning for Photonic Component Design
by: Witt, Donald, et al.
Published: (2023)
by: Witt, Donald, et al.
Published: (2023)
FedKL: Tackling Data Heterogeneity in Federated Reinforcement Learning by Penalizing KL Divergence
by: Xie, Zhijie, et al.
Published: (2022)
by: Xie, Zhijie, et al.
Published: (2022)
diff History for Neural Language Agents
by: Piterbarg, Ulyana, et al.
Published: (2023)
by: Piterbarg, Ulyana, et al.
Published: (2023)
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
by: Saporta, Adriel, et al.
Published: (2024)
by: Saporta, Adriel, et al.
Published: (2024)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Collapsing Taylor Mode Automatic Differentiation
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
Three Forms of Stochastic Injection for Improved Distribution-to-Distribution Generative Modeling
by: Su, Shiye, et al.
Published: (2025)
by: Su, Shiye, et al.
Published: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024)
by: Zhao, Heyang, et al.
Published: (2024)
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
by: Chen, Claire, et al.
Published: (2026)
by: Chen, Claire, et al.
Published: (2026)
Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
by: Sinha, Abhijeet, et al.
Published: (2026)
by: Sinha, Abhijeet, et al.
Published: (2026)
Mode Collapse of Mean-Field Variational Inference
by: Sheng, Shunan, et al.
Published: (2025)
by: Sheng, Shunan, et al.
Published: (2025)
Stochastic interpolants with data-dependent couplings
by: Albergo, Michael S., et al.
Published: (2023)
by: Albergo, Michael S., et al.
Published: (2023)
Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
by: Shan, Zhao, et al.
Published: (2024)
by: Shan, Zhao, et al.
Published: (2024)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
by: Yu, Xin, et al.
Published: (2026)
by: Yu, Xin, et al.
Published: (2026)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Hard-Negative Sampling for Contrastive Learning: Optimal Representation Geometry and Neural- vs Dimensional-Collapse
by: Jiang, Ruijie, et al.
Published: (2023)
by: Jiang, Ruijie, et al.
Published: (2023)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
by: Liu, Kezhao, et al.
Published: (2025)
by: Liu, Kezhao, et al.
Published: (2025)
Mildly Conservative Regularized Evaluation for Offline Reinforcement Learning
by: Chen, Haohui, et al.
Published: (2025)
by: Chen, Haohui, et al.
Published: (2025)
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization
by: He, Junlin, et al.
Published: (2024)
by: He, Junlin, et al.
Published: (2024)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
by: Prakash, Hari K, et al.
Published: (2026)
by: Prakash, Hari K, et al.
Published: (2026)
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022)
by: Puli, Aahlad, et al.
Published: (2022)
Similar Items
-
Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction
by: GX-Chen, Anthony, et al.
Published: (2024) -
Attention and Compression is all you need for Controllably Efficient Language Models
by: Prakash, Jatin, et al.
Published: (2025) -
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
by: Jhaveri, Ayush Rajesh, et al.
Published: (2026) -
Light-weight probing of unsupervised representations for Reinforcement Learning
by: Zhang, Wancong, et al.
Published: (2022) -
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
by: Dragutinović, Sara, et al.
Published: (2026)