On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Rishabh, Vieillard, Nino, Zhou, Yongchao, Stanczyk, Piotr, Ramos, Sabela, Geist, Matthieu, Bachem, Olivier
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909075367788544
author Agarwal, Rishabh
Vieillard, Nino
Zhou, Yongchao
Stanczyk, Piotr
Ramos, Sabela
Geist, Matthieu
Bachem, Olivier
author_facet Agarwal, Rishabh
Vieillard, Nino
Zhou, Yongchao
Stanczyk, Piotr
Ramos, Sabela
Geist, Matthieu
Bachem, Olivier
contents Knowledge distillation (KD) is widely used for compressing a teacher model to reduce its inference cost and memory footprint, by training a smaller student model. However, current KD methods for auto-regressive sequence models suffer from distribution mismatch between output sequences seen during training and those generated by the student during inference. To address this issue, we introduce Generalized Knowledge Distillation (GKD). Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-generated output sequences by leveraging feedback from the teacher on such sequences. Unlike supervised KD approaches, GKD also offers the flexibility to employ alternative loss functions between the student and teacher, which can be useful when the student lacks the expressivity to mimic the teacher's distribution. Furthermore, GKD facilitates the seamless integration of distillation with RL fine-tuning (RLHF). We demonstrate the efficacy of GKD for distilling auto-regressive language models on summarization, translation, and arithmetic reasoning tasks, and task-agnostic distillation for instruction-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2306_13649
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Agarwal, Rishabh
Vieillard, Nino
Zhou, Yongchao
Stanczyk, Piotr
Ramos, Sabela
Geist, Matthieu
Bachem, Olivier
Machine Learning
Artificial Intelligence
Computation and Language
Knowledge distillation (KD) is widely used for compressing a teacher model to reduce its inference cost and memory footprint, by training a smaller student model. However, current KD methods for auto-regressive sequence models suffer from distribution mismatch between output sequences seen during training and those generated by the student during inference. To address this issue, we introduce Generalized Knowledge Distillation (GKD). Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-generated output sequences by leveraging feedback from the teacher on such sequences. Unlike supervised KD approaches, GKD also offers the flexibility to employ alternative loss functions between the student and teacher, which can be useful when the student lacks the expressivity to mimic the teacher's distribution. Furthermore, GKD facilitates the seamless integration of distillation with RL fine-tuning (RLHF). We demonstrate the efficacy of GKD for distilling auto-regressive language models on summarization, translation, and arithmetic reasoning tasks, and task-agnostic distillation for instruction-tuning.
title On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2306.13649