A Comedy of Estimators: On KL Regularization in RL Training of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Shah, Vedant, Obando-Ceron, Johan, Jain, Vineet, Bartoldson, Brian, Kailkhura, Bhavya, Mittal, Sarthak, Berseth, Glen, Castro, Pablo Samuel, Bengio, Yoshua, Malkin, Nikolay, Jain, Moksh, Venkatraman, Siddarth, Courville, Aaron |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
por: Venkatraman, Siddarth, et al.
Publicado: (2025)
por: Venkatraman, Siddarth, et al.
Publicado: (2025)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
por: Bartoldson, Brian, et al.
Publicado: (2025)
por: Bartoldson, Brian, et al.
Publicado: (2025)
Solving Bayesian inverse problems with diffusion priors and off-policy RL
por: Scimeca, Luca, et al.
Publicado: (2025)
por: Scimeca, Luca, et al.
Publicado: (2025)
Amortizing intractable inference in diffusion models for vision, language, and control
por: Venkatraman, Siddarth, et al.
Publicado: (2024)
por: Venkatraman, Siddarth, et al.
Publicado: (2024)
Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
por: Venkatraman, Siddarth, et al.
Publicado: (2025)
por: Venkatraman, Siddarth, et al.
Publicado: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
por: Mittal, Sarthak, et al.
Publicado: (2025)
por: Mittal, Sarthak, et al.
Publicado: (2025)
Machine learning and information theory concepts towards an AI Mathematician
por: Bengio, Yoshua, et al.
Publicado: (2024)
por: Bengio, Yoshua, et al.
Publicado: (2024)
Amortizing intractable inference in large language models
por: Hu, Edward J., et al.
Publicado: (2023)
por: Hu, Edward J., et al.
Publicado: (2023)
Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
por: Tang, Hongyao, et al.
Publicado: (2025)
por: Tang, Hongyao, et al.
Publicado: (2025)
PhyloGFN: Phylogenetic inference with generative flow networks
por: Zhou, Mingyang, et al.
Publicado: (2023)
por: Zhou, Mingyang, et al.
Publicado: (2023)
Action abstractions for amortized sampling
por: Boussif, Oussama, et al.
Publicado: (2024)
por: Boussif, Oussama, et al.
Publicado: (2024)
On Generalization for Generative Flow Networks
por: Krichel, Anas, et al.
Publicado: (2024)
por: Krichel, Anas, et al.
Publicado: (2024)
Multi-Fidelity Active Learning with GFlowNets
por: Hernandez-Garcia, Alex, et al.
Publicado: (2023)
por: Hernandez-Garcia, Alex, et al.
Publicado: (2023)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
por: Wang, Zeyu, et al.
Publicado: (2025)
por: Wang, Zeyu, et al.
Publicado: (2025)
Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
por: Castanyer, Roger Creus, et al.
Publicado: (2025)
por: Castanyer, Roger Creus, et al.
Publicado: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
por: Bartoldson, Brian R., et al.
Publicado: (2024)
por: Bartoldson, Brian R., et al.
Publicado: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
por: Lee, Seanie, et al.
Publicado: (2024)
por: Lee, Seanie, et al.
Publicado: (2024)
Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning
por: Ho, Matthew, et al.
Publicado: (2024)
por: Ho, Matthew, et al.
Publicado: (2024)
Improved off-policy training of diffusion samplers
por: Sendera, Marcin, et al.
Publicado: (2024)
por: Sendera, Marcin, et al.
Publicado: (2024)
Discrete Probabilistic Inference as Control in Multi-path Environments
por: Deleu, Tristan, et al.
Publicado: (2024)
por: Deleu, Tristan, et al.
Publicado: (2024)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
por: Sokar, Ghada, et al.
Publicado: (2024)
por: Sokar, Ghada, et al.
Publicado: (2024)
In value-based deep reinforcement learning, a pruned network is a good network
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
Neuroplastic Expansion in Deep Reinforcement Learning
por: Liu, Jiashun, et al.
Publicado: (2024)
por: Liu, Jiashun, et al.
Publicado: (2024)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
por: McDonald, Tavish, et al.
Publicado: (2025)
por: McDonald, Tavish, et al.
Publicado: (2025)
Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
por: Berseth, Glen
Publicado: (2025)
por: Berseth, Glen
Publicado: (2025)
Towards DNA-Encoded Library Generation with GFlowNets
por: Koziarski, Michał, et al.
Publicado: (2024)
por: Koziarski, Michał, et al.
Publicado: (2024)
Adaptive Computation Pruning for the Forgetting Transformer
por: Lin, Zhixuan, et al.
Publicado: (2025)
por: Lin, Zhixuan, et al.
Publicado: (2025)
The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
por: Mayor, Walter, et al.
Publicado: (2025)
por: Mayor, Walter, et al.
Publicado: (2025)
Book Review
por: Bhavya Jain
Publicado: (2024)
por: Bhavya Jain
Publicado: (2024)
Discrete, compositional, and symbolic representations through attractor dynamics
por: Nam, Andrew, et al.
Publicado: (2023)
por: Nam, Andrew, et al.
Publicado: (2023)
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
por: Kim, Minsu, et al.
Publicado: (2025)
por: Kim, Minsu, et al.
Publicado: (2025)
Intelligent Switching for Reset-Free RL
por: Patil, Darshan, et al.
Publicado: (2024)
por: Patil, Darshan, et al.
Publicado: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
por: Geiping, Jonas, et al.
Publicado: (2025)
por: Geiping, Jonas, et al.
Publicado: (2025)
Expected flow networks in stochastic environments and two-player zero-sum games
por: Jiralerspong, Marco, et al.
Publicado: (2023)
por: Jiralerspong, Marco, et al.
Publicado: (2023)
Improving Robustness In Sparse Autoencoders via Masked Regularization
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
Distributional GFlowNets with Quantile Flows
por: Zhang, Dinghuai, et al.
Publicado: (2023)
por: Zhang, Dinghuai, et al.
Publicado: (2023)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
por: Liu, Jiashun, et al.
Publicado: (2025)
por: Liu, Jiashun, et al.
Publicado: (2025)
Learning Decision Trees as Amortized Structure Inference
por: Mahfoud, Mohammed, et al.
Publicado: (2025)
por: Mahfoud, Mohammed, et al.
Publicado: (2025)
Efficient Causal Graph Discovery Using Large Language Models
por: Jiralerspong, Thomas, et al.
Publicado: (2024)
por: Jiralerspong, Thomas, et al.
Publicado: (2024)
Ejemplares similares
-
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
por: Venkatraman, Siddarth, et al.
Publicado: (2025) -
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
por: Bartoldson, Brian, et al.
Publicado: (2025) -
Solving Bayesian inverse problems with diffusion priors and off-policy RL
por: Scimeca, Luca, et al.
Publicado: (2025) -
Amortizing intractable inference in diffusion models for vision, language, and control
por: Venkatraman, Siddarth, et al.
Publicado: (2024) -
Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
por: Venkatraman, Siddarth, et al.
Publicado: (2025)