Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Arnal, Charles, Narozniak, Gaëtan, Cabannes, Vivien, Tang, Yunhao, Kempe, Julia, Munos, Remi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Outcome-based Exploration for LLM Reasoning
by: Song, Yuda, et al.
Published: (2025)
by: Song, Yuda, et al.
Published: (2025)
Touring sampling with pushforward maps
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
Learning with Hidden Factorial Structure
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Provable Benefits of In-Tool Learning for Large Language Models
by: Houliston, Sam, et al.
Published: (2025)
by: Houliston, Sam, et al.
Published: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Formalizing Mathematics at Scale
by: Rammal, Ahmad, et al.
Published: (2026)
by: Rammal, Ahmad, et al.
Published: (2026)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
by: Rowland, Mark, et al.
Published: (2024)
by: Rowland, Mark, et al.
Published: (2024)
VA-learning as a more efficient alternative to Q-learning
by: Tang, Yunhao, et al.
Published: (2023)
by: Tang, Yunhao, et al.
Published: (2023)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
by: Jiang, Guochao, et al.
Published: (2026)
by: Jiang, Guochao, et al.
Published: (2026)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
by: Yao, Chaorui, et al.
Published: (2025)
by: Yao, Chaorui, et al.
Published: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Clustering Head: A Visual Case Study of the Training Dynamics in Transformers
by: Odonnat, Ambroise, et al.
Published: (2024)
by: Odonnat, Ambroise, et al.
Published: (2024)
Automatic Textbook Formalization
by: Gloeckle, Fabian, et al.
Published: (2026)
by: Gloeckle, Fabian, et al.
Published: (2026)
Distilling LLM Feedback for Lean Theorem Proving
by: Narozniak, Gaetan, et al.
Published: (2026)
by: Narozniak, Gaetan, et al.
Published: (2026)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
Easing Optimization Paths: a Circuit Perspective
by: Odonnat, Ambroise, et al.
Published: (2025)
by: Odonnat, Ambroise, et al.
Published: (2025)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
by: Xia, Yinan, et al.
Published: (2026)
by: Xia, Yinan, et al.
Published: (2026)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Super-Exponential Regret for UCT, AlphaGo and Variants
by: Orseau, Laurent, et al.
Published: (2024)
by: Orseau, Laurent, et al.
Published: (2024)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
by: Hou, Zhenyu, et al.
Published: (2025)
by: Hou, Zhenyu, et al.
Published: (2025)
LIRE: listwise reward enhancement for preference alignment
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025)
by: Tsilivis, Nikolaos, et al.
Published: (2025)
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
by: Roux, Nicolas Le, et al.
Published: (2025)
by: Roux, Nicolas Le, et al.
Published: (2025)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
by: Sundaram, Shobhita, et al.
Published: (2026)
by: Sundaram, Shobhita, et al.
Published: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
Similar Items
-
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026) -
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024) -
Outcome-based Exploration for LLM Reasoning
by: Song, Yuda, et al.
Published: (2025) -
Touring sampling with pushforward maps
by: Cabannes, Vivien, et al.
Published: (2023) -
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)