Self-Distillation Enables Continual Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shenfeld, Idan, Damani, Mehul, Hübotter, Jonas, Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023)
by: Shenfeld, Idan, et al.
Published: (2023)
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
by: Bahlous-Boldi, Ryan, et al.
Published: (2026)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Aligning Language Models from User Interactions
by: Buening, Thomas Kleine, et al.
Published: (2026)
by: Buening, Thomas Kleine, et al.
Published: (2026)
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)
by: Krause, Andreas, et al.
Published: (2025)
Curiosity-driven Red-teaming for Large Language Models
by: Hong, Zhang-Wei, et al.
Published: (2024)
by: Hong, Zhang-Wei, et al.
Published: (2024)
CoordLight: Learning Decentralized Coordination for Network-Wide Traffic Signal Control
by: Zhang, Yifeng, et al.
Published: (2026)
by: Zhang, Yifeng, et al.
Published: (2026)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
by: Acikgoz, Emre Can, et al.
Published: (2026)
by: Acikgoz, Emre Can, et al.
Published: (2026)
Test-time Offline Reinforcement Learning on Goal-related Experience
by: Bagatella, Marco, et al.
Published: (2025)
by: Bagatella, Marco, et al.
Published: (2025)
Transductive Active Learning: Theory and Applications
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
by: Otth, Matthias, et al.
Published: (2025)
by: Otth, Matthias, et al.
Published: (2025)
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Active Fine-Tuning of Multi-Task Policies
by: Bagatella, Marco, et al.
Published: (2024)
by: Bagatella, Marco, et al.
Published: (2024)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Self-Adapting Language Models
by: Zweiger, Adam, et al.
Published: (2025)
by: Zweiger, Adam, et al.
Published: (2025)
LITE: Efficiently Estimating Gaussian Probability of Maximality
by: Menet, Nicolas, et al.
Published: (2025)
by: Menet, Nicolas, et al.
Published: (2025)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025)
by: Bertolissi, Ryo, et al.
Published: (2025)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
H2LooP Spark Preview: Continual Pretraining of Large Language Models for Low-Level Embedded Systems Code
by: Singh, Amit, et al.
Published: (2026)
by: Singh, Amit, et al.
Published: (2026)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024)
by: Reuss, Moritz, et al.
Published: (2024)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Majority Voting for Code Generation
by: Launer, Tim, et al.
Published: (2026)
by: Launer, Tim, et al.
Published: (2026)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
Random Latent Exploration for Deep Reinforcement Learning
by: Mahankali, Srinath, et al.
Published: (2024)
by: Mahankali, Srinath, et al.
Published: (2024)
Efficient privacy loss accounting for subsampling and random allocation
by: Feldman, Vitaly, et al.
Published: (2026)
by: Feldman, Vitaly, et al.
Published: (2026)
Privacy amplification by random allocation
by: Feldman, Vitaly, et al.
Published: (2025)
by: Feldman, Vitaly, et al.
Published: (2025)
Similar Items
-
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025) -
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024) -
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024) -
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023) -
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)