RL's Razor: Why Online Reinforcement Learning Forgets Less
Fuente:
arXiv
Saved in:
| Main Authors: | Shenfeld, Idan, Pari, Jyothish, Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024)
by: Reuss, Moritz, et al.
Published: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023)
by: Shenfeld, Idan, et al.
Published: (2023)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
General Intelligence Requires Reward-based Pretraining
by: Han, Seungwook, et al.
Published: (2025)
by: Han, Seungwook, et al.
Published: (2025)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Self-Adapting Language Models
by: Zweiger, Adam, et al.
Published: (2025)
by: Zweiger, Adam, et al.
Published: (2025)
Few-Shot Task Learning through Inverse Generative Modeling
by: Netanyahu, Aviv, et al.
Published: (2024)
by: Netanyahu, Aviv, et al.
Published: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Automatic Environment Shaping is the Next Frontier in RL
by: Park, Younghyo, et al.
Published: (2024)
by: Park, Younghyo, et al.
Published: (2024)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
Curiosity-driven Red-teaming for Large Language Models
by: Hong, Zhang-Wei, et al.
Published: (2024)
by: Hong, Zhang-Wei, et al.
Published: (2024)
LoRA Learns Less and Forgets Less
by: Biderman, Dan, et al.
Published: (2024)
by: Biderman, Dan, et al.
Published: (2024)
Random Latent Exploration for Deep Reinforcement Learning
by: Mahankali, Srinath, et al.
Published: (2024)
by: Mahankali, Srinath, et al.
Published: (2024)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Why Zeroth-Order Adaptation May Forget Less: A Randomized Shaping Theory
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Leveraging Manifold Embeddings for Enhanced Graph Transformer Representations and Learning
by: Jyothish, Ankit, et al.
Published: (2025)
by: Jyothish, Ankit, et al.
Published: (2025)
A Geometric Modeling of Occam's Razor in Deep Learning
by: Sun, Ke, et al.
Published: (2019)
by: Sun, Ke, et al.
Published: (2019)
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
by: Bagaria, Vaidehi, et al.
Published: (2026)
by: Bagaria, Vaidehi, et al.
Published: (2026)
Online Learning of Neural Networks
by: Daniely, Amit, et al.
Published: (2025)
by: Daniely, Amit, et al.
Published: (2025)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Catastrophic Forgetting Mitigation Through Plateau Phase Activity Profiling
by: Mashiach, Idan, et al.
Published: (2025)
by: Mashiach, Idan, et al.
Published: (2025)
Learning from Less: SINDy Surrogates in RL
by: Dixit, Aniket, et al.
Published: (2025)
by: Dixit, Aniket, et al.
Published: (2025)
Forget Less by Learning Together through Concept Consolidation
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
by: Sabbaghi, Mahdi, et al.
Published: (2026)
by: Sabbaghi, Mahdi, et al.
Published: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning
by: Attias, Idan, et al.
Published: (2026)
by: Attias, Idan, et al.
Published: (2026)
Occam's Razor is Only as Sharp as Your ELBO
by: Harvey, Ethan, et al.
Published: (2026)
by: Harvey, Ethan, et al.
Published: (2026)
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024)
by: Schulte, Oliver, et al.
Published: (2024)
Forget Less by Learning from Parents Through Hierarchical Relationships
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
by: Doron-Arad, Ilan, et al.
Published: (2026)
by: Doron-Arad, Ilan, et al.
Published: (2026)
Tradeoffs between Mistakes and ERM Oracle Calls in Online and Transductive Online Learning
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Similar Items
-
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024) -
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024) -
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023) -
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024) -
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)