Knowledge Distillation with Training Wheels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Guanlin, Ramachandran, Anand, Gangwani, Tanmay, Fu, Yan, Sethy, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909509496078336
author Liu, Guanlin
Ramachandran, Anand
Gangwani, Tanmay
Fu, Yan
Sethy, Abhinav
author_facet Liu, Guanlin
Ramachandran, Anand
Gangwani, Tanmay
Fu, Yan
Sethy, Abhinav
contents Knowledge distillation is used, in generative language modeling, to train a smaller student model using the help of a larger teacher model, resulting in improved capabilities for the student model. In this paper, we formulate a more general framework for knowledge distillation where the student learns from the teacher during training, and also learns to ask for the teacher's help at test-time following rules specifying test-time restrictions. Towards this, we first formulate knowledge distillation as an entropy-regularized value optimization problem. Adopting Path Consistency Learning to solve this, leads to a new knowledge distillation algorithm using on-policy and off-policy demonstrations. We extend this using constrained reinforcement learning to a framework that incorporates the use of the teacher model as a test-time reference, within constraints. In this situation, akin to a human learner, the model needs to learn not only the learning material, but also the relative difficulty of different sections to prioritize for seeking teacher help. We examine the efficacy of our method through experiments in translation and summarization tasks, observing trends in accuracy and teacher use, noting that our approach unlocks operating points not available to the popular Speculative Decoding approach.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge Distillation with Training Wheels
Liu, Guanlin
Ramachandran, Anand
Gangwani, Tanmay
Fu, Yan
Sethy, Abhinav
Computation and Language
Machine Learning
Knowledge distillation is used, in generative language modeling, to train a smaller student model using the help of a larger teacher model, resulting in improved capabilities for the student model. In this paper, we formulate a more general framework for knowledge distillation where the student learns from the teacher during training, and also learns to ask for the teacher's help at test-time following rules specifying test-time restrictions. Towards this, we first formulate knowledge distillation as an entropy-regularized value optimization problem. Adopting Path Consistency Learning to solve this, leads to a new knowledge distillation algorithm using on-policy and off-policy demonstrations. We extend this using constrained reinforcement learning to a framework that incorporates the use of the teacher model as a test-time reference, within constraints. In this situation, akin to a human learner, the model needs to learn not only the learning material, but also the relative difficulty of different sections to prioritize for seeking teacher help. We examine the efficacy of our method through experiments in translation and summarization tasks, observing trends in accuracy and teacher use, noting that our approach unlocks operating points not available to the popular Speculative Decoding approach.
title Knowledge Distillation with Training Wheels
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.17717