Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Roux, Nicolas Le, Bellemare, Marc G., Lebensold, Jonathan, Bergeron, Arnaud, Greaves, Joshua, Fréchette, Alex, Pelletier, Carolyne, Thibodeau-Laufer, Eric, Toth, Sándor, Work, Sam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
by: Yao, Chaorui, et al.
Published: (2025)
by: Yao, Chaorui, et al.
Published: (2025)
Taper-based scattering formulation of the Helmholtz equation to improve the training process of Physics-Informed Neural Networks
by: Dörfler, W., et al.
Published: (2024)
by: Dörfler, W., et al.
Published: (2024)
Research Directions for Verifiable Crypto-Physically Secure TEEs
by: Bellemare, Sylvain
Published: (2024)
by: Bellemare, Sylvain
Published: (2024)
Marc Bellemare
by: Marc Bellemare
Published: (2024)
by: Marc Bellemare
Published: (2024)
QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE
by: Zhao, Junjie, et al.
Published: (2024)
by: Zhao, Junjie, et al.
Published: (2024)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
by: Han, Haoyu, et al.
Published: (2026)
by: Han, Haoyu, et al.
Published: (2026)
THE CONVENTION: THE SOLUTION TO REINFORCE ESDP?
by: MANUEL VÁZQUEZ MUÑOZ
Published: (2003)
by: MANUEL VÁZQUEZ MUÑOZ
Published: (2003)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Stable oxygen isotope analysis of water samples during helicopter/ice camp TRANSDRIFT-XX, Laptev Sea
by: Bauch, Dorothea, et al.
Published: (2020)
by: Bauch, Dorothea, et al.
Published: (2020)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
by: Ahmadian, Arash, et al.
Published: (2024)
by: Ahmadian, Arash, et al.
Published: (2024)
PODCASTING: TOOL TO DEVELOP AND REINFORCE EFL LEARNING
by: Anali Carolina Rodríguez-Castro
Published: (2023)
by: Anali Carolina Rodríguez-Castro
Published: (2023)
TH-1517-DFD: Autonomous Interstellar Reconnaissance Spacecraft - Complete System Architecture & Mond Self-Repair
by: Thibodeau, Pascal
Published: (2026)
by: Thibodeau, Pascal
Published: (2026)
L'Humanité a Choisi la Mauvaise Guerre depuis 5 500 ans : Analyse métrologique ThibEquation v6.1 du vecteur guerrier et de la réponse institutionnelle à 3I/ATLAS
by: Thibodeau, Pascal
Published: (2026)
by: Thibodeau, Pascal
Published: (2026)
On the Privacy of Selection Mechanisms with Gaussian Noise
by: Lebensold, Jonathan, et al.
Published: (2024)
by: Lebensold, Jonathan, et al.
Published: (2024)
REINFORCE-ING Chemical Language Models for Drug Discovery
by: Thomas, Morgan, et al.
Published: (2025)
by: Thomas, Morgan, et al.
Published: (2025)
Inclusive Science in Canadian Federal Science Laboratories: Exploring Accessibility Policies and Best Practices
by: Sacha Ghandeharian, et al.
Published: (2026)
by: Sacha Ghandeharian, et al.
Published: (2026)
EVENTOS EXTREMOS DE PRECIPITAÇÃO NO ESTADO DO PARANÁ
by: Carolyne B. MACHADO
Published: (2013)
by: Carolyne B. MACHADO
Published: (2013)
Anatomía y fisiología / Gary A. Thibodeau, Kevin T. Patton; traductor, Diorki Servicios Integrales de Edición
by: Thibodeau, Gary A
Published: (2007)
by: Thibodeau, Gary A
Published: (2007)
Global agricultural value chains and food prices
by: Bernhard Dalheimer, et al.
Published: (2025)
by: Bernhard Dalheimer, et al.
Published: (2025)
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
by: Jian, Hu
Published: (2025)
by: Jian, Hu
Published: (2025)
Complexity, Features, and Comparisons in Forensic Handwriting Examination
by: Kylie Jones, et al.
Published: (2024)
by: Kylie Jones, et al.
Published: (2024)
PROJETO DE EXTENSÃO ILUMINE: A ENTRADA DA FIGURA DO PALHAÇO NO AMBIENTE HOSPITALAR
by: Cely Carolyne Pontes MORCERF
Published: (2015)
by: Cely Carolyne Pontes MORCERF
Published: (2015)
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
by: Geisler, Simon, et al.
Published: (2025)
by: Geisler, Simon, et al.
Published: (2025)
Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
by: Yu, Peiyu, et al.
Published: (2025)
by: Yu, Peiyu, et al.
Published: (2025)
Galahs near Melbourne
by: Greaves, T.
Published: (1928)
by: Greaves, T.
Published: (1928)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
Ultra-Wideband Tapered Transducers in Thin-Film Lithium Niobate on Silicon Carbide
by: Kramer, Jack, et al.
Published: (2024)
by: Kramer, Jack, et al.
Published: (2024)
Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
by: Rahn, Nate, et al.
Published: (2023)
by: Rahn, Nate, et al.
Published: (2023)
Comp-LTL: Temporal Logic Planning via Zero-Shot Policy Composition
by: Bergeron, Taylor, et al.
Published: (2024)
by: Bergeron, Taylor, et al.
Published: (2024)
Non‐Original and Digitally Captured Handwriting: Considerations for Forensic Handwriting Examinations
by: Kylie Jones, et al.
Published: (2024)
by: Kylie Jones, et al.
Published: (2024)
Predicting Flare in Patients With Rheumatoid Arthritis in Biologic Induced Remission, on Tapering, and on Stable Therapy
by: Hanna Gul, et al.
Published: (2024)
by: Hanna Gul, et al.
Published: (2024)
Recognizing the legitimacy of a deep unease: improving the analysis of systemic discrimination by considering microaggressions experienced in a society in denial about racism and sexism
by: Karine Bellemare (Author), et al.
Published: (2021)
by: Karine Bellemare (Author), et al.
Published: (2021)
Geometry of efficient weight vectors
by: Ábele-Nagy, Kristóf, et al.
Published: (2025)
by: Ábele-Nagy, Kristóf, et al.
Published: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Superpositions of tree-tensor networks for single-reference ground states in the strong correlation regime
by: Bergeron, Dominic
Published: (2024)
by: Bergeron, Dominic
Published: (2024)
El desarrollo psicológico del niño : desde la primera edad hasta la adolescencia / M. Bergeron; traductor, G. Gonzalvo Mainar
by: Bergeron, M
Published: (2000)
by: Bergeron, M
Published: (2000)
La epoca de las revoluciones europeas : 1780-1848 / Louis Bergeron, Francois Furet, Reinhart Koselleck ; traducción de Francisco Pérez Gutierrez
by: Bergeron, Louis
by: Bergeron, Louis
A Qualitative Case Study Approach To Examine Information Resources Management. (Utilisation d'une Approche Qualitative par Methode de cas pour Etudier la Gestion des Ressources D'information).
by: Bergeron, Pierrette
Published: (1997)
by: Bergeron, Pierrette
Published: (1997)
Similar Items
-
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
by: Yao, Chaorui, et al.
Published: (2025) -
Taper-based scattering formulation of the Helmholtz equation to improve the training process of Physics-Informed Neural Networks
by: Dörfler, W., et al.
Published: (2024) -
Research Directions for Verifiable Crypto-Physically Secure TEEs
by: Bellemare, Sylvain
Published: (2024) -
Marc Bellemare
by: Marc Bellemare
Published: (2024) -
QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE
by: Zhao, Junjie, et al.
Published: (2024)