Self-Distillation Enables Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shenfeld, Idan, Damani, Mehul, Hübotter, Jonas, Agrawal, Pulkit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910002556436480
author Shenfeld, Idan
Damani, Mehul
Hübotter, Jonas
Agrawal, Pulkit
author_facet Shenfeld, Idan
Damani, Mehul
Hübotter, Jonas
Agrawal, Pulkit
contents Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. SDFT leverages in-context learning by using a demonstration-conditioned model as its own teacher, generating on-policy training signals that preserve prior capabilities while acquiring new skills. Across skill learning and knowledge acquisition tasks, SDFT consistently outperforms SFT, achieving higher new-task accuracy while substantially reducing catastrophic forgetting. In sequential learning experiments, SDFT enables a single model to accumulate multiple skills over time without performance regression, establishing on-policy distillation as a practical path to continual learning from demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19897
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Distillation Enables Continual Learning
Shenfeld, Idan
Damani, Mehul
Hübotter, Jonas
Agrawal, Pulkit
Machine Learning
Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explicit reward functions that are often unavailable. Learning from expert demonstrations, the primary alternative, is dominated by supervised fine-tuning (SFT), which is inherently off-policy. We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. SDFT leverages in-context learning by using a demonstration-conditioned model as its own teacher, generating on-policy training signals that preserve prior capabilities while acquiring new skills. Across skill learning and knowledge acquisition tasks, SDFT consistently outperforms SFT, achieving higher new-task accuracy while substantially reducing catastrophic forgetting. In sequential learning experiments, SDFT enables a single model to accumulate multiple skills over time without performance regression, establishing on-policy distillation as a practical path to continual learning from demonstrations.
title Self-Distillation Enables Continual Learning
topic Machine Learning
url https://arxiv.org/abs/2601.19897