Don't Throw Away Data: Better Sequence Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jun, Briakou, Eleftheria, Dadkhahi, Hamid, Agarwal, Rishabh, Cherry, Colin, Cohn, Trevor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913430785490944
author Wang, Jun
Briakou, Eleftheria
Dadkhahi, Hamid
Agarwal, Rishabh
Cherry, Colin
Cohn, Trevor
author_facet Wang, Jun
Briakou, Eleftheria
Dadkhahi, Hamid
Agarwal, Rishabh
Cherry, Colin
Cohn, Trevor
contents A critical component in knowledge distillation is the means of coupling the teacher and student. The predominant sequence knowledge distillation method involves supervised learning of the student against teacher-decoded outputs, and is exemplified by the current state of the art, which incorporates minimum Bayes risk (MBR) decoding. In this paper we seek to integrate MBR more tightly in distillation training, specifically by using several high scoring MBR translations, rather than a single selected sequence, thus capturing a rich diversity of teacher outputs. Our experiments on English to German and English to Japanese translation show consistent improvements over strong baseline methods for both tasks and with varying model sizes. Additionally, we conduct a detailed analysis focusing on data efficiency and capacity curse aspects to elucidate MBR-n and explore its further potential.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10456
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Don't Throw Away Data: Better Sequence Knowledge Distillation
Wang, Jun
Briakou, Eleftheria
Dadkhahi, Hamid
Agarwal, Rishabh
Cherry, Colin
Cohn, Trevor
Computation and Language
A critical component in knowledge distillation is the means of coupling the teacher and student. The predominant sequence knowledge distillation method involves supervised learning of the student against teacher-decoded outputs, and is exemplified by the current state of the art, which incorporates minimum Bayes risk (MBR) decoding. In this paper we seek to integrate MBR more tightly in distillation training, specifically by using several high scoring MBR translations, rather than a single selected sequence, thus capturing a rich diversity of teacher outputs. Our experiments on English to German and English to Japanese translation show consistent improvements over strong baseline methods for both tasks and with varying model sizes. Additionally, we conduct a detailed analysis focusing on data efficiency and capacity curse aspects to elucidate MBR-n and explore its further potential.
title Don't Throw Away Data: Better Sequence Knowledge Distillation
topic Computation and Language
url https://arxiv.org/abs/2407.10456