AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911527243612160 |
|---|---|
| author | Peng, Jingyu Wang, Maolin Cai, Hengyi Li, Yuchen Zhang, Kai Wang, Shuaiqiang Yin, Dawei Zhao, Xiangyu |
| author_facet | Peng, Jingyu Wang, Maolin Cai, Hengyi Li, Yuchen Zhang, Kai Wang, Shuaiqiang Yin, Dawei Zhao, Xiangyu |
| contents | Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_07842 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching Peng, Jingyu Wang, Maolin Cai, Hengyi Li, Yuchen Zhang, Kai Wang, Shuaiqiang Yin, Dawei Zhao, Xiangyu Computation and Language Artificial Intelligence Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead. |
| title | AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.07842 |