AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Peng, Jingyu, Wang, Maolin, Cai, Hengyi, Li, Yuchen, Zhang, Kai, Wang, Shuaiqiang, Yin, Dawei, Zhao, Xiangyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911527243612160
author Peng, Jingyu
Wang, Maolin
Cai, Hengyi
Li, Yuchen
Zhang, Kai
Wang, Shuaiqiang
Yin, Dawei
Zhao, Xiangyu
author_facet Peng, Jingyu
Wang, Maolin
Cai, Hengyi
Li, Yuchen
Zhang, Kai
Wang, Shuaiqiang
Yin, Dawei
Zhao, Xiangyu
contents Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
Peng, Jingyu
Wang, Maolin
Cai, Hengyi
Li, Yuchen
Zhang, Kai
Wang, Shuaiqiang
Yin, Dawei
Zhao, Xiangyu
Computation and Language
Artificial Intelligence
Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead.
title AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.07842