Born Again Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Furlanello, Tommaso, Lipton, Zachary C., Tschannen, Michael, Itti, Laurent, Anandkumar, Anima
Formato: Preprint
Publicado: 2018
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910351307571200
author Furlanello, Tommaso
Lipton, Zachary C.
Tschannen, Michael
Itti, Laurent
Anandkumar, Anima
author_facet Furlanello, Tommaso
Lipton, Zachary C.
Tschannen, Michael
Itti, Laurent
Anandkumar, Anima
contents Knowledge Distillation (KD) consists of transferring “knowledge” from one machine learning model (the teacher) to another (the student). Commonly, the teacher is a high-capacity model with formidable performance, while the student is more compact. By transferring knowledge, one hopes to benefit from the student’s compactness, without sacrificing too much performance. We study KD from a new perspective: rather than compressing models, we train students parameterized identically to their teachers. Surprisingly, these Born-Again Networks (BANs), outperform their teachers significantly, both on computer vision and language modeling tasks. Our experiments with BANs based on DenseNets demonstrate state-of-the-art performance on the CIFAR-10 (3.5%) and CIFAR-100 (15.5%) datasets, by validation error. Additional experiments explore two distillation objectives: (i) Confidence-Weighted by Teacher Max (CWTM) and (ii) Dark Knowledge with Permuted Predictions (DKPP). Both methods elucidate the essential components of KD, demonstrating the effect of the teacher outputs on both predicted and non-predicted classes.
format Preprint
id arxiv_https___arxiv_org_abs_1805_04770
institution arXiv
publishDate 2018
record_format arxiv
spellingShingle Born Again Neural Networks
Furlanello, Tommaso
Lipton, Zachary C.
Tschannen, Michael
Itti, Laurent
Anandkumar, Anima
Machine Learning
Artificial Intelligence
Knowledge Distillation (KD) consists of transferring “knowledge” from one machine learning model (the teacher) to another (the student). Commonly, the teacher is a high-capacity model with formidable performance, while the student is more compact. By transferring knowledge, one hopes to benefit from the student’s compactness, without sacrificing too much performance. We study KD from a new perspective: rather than compressing models, we train students parameterized identically to their teachers. Surprisingly, these Born-Again Networks (BANs), outperform their teachers significantly, both on computer vision and language modeling tasks. Our experiments with BANs based on DenseNets demonstrate state-of-the-art performance on the CIFAR-10 (3.5%) and CIFAR-100 (15.5%) datasets, by validation error. Additional experiments explore two distillation objectives: (i) Confidence-Weighted by Teacher Max (CWTM) and (ii) Dark Knowledge with Permuted Predictions (DKPP). Both methods elucidate the essential components of KD, demonstrating the effect of the teacher outputs on both predicted and non-predicted classes.
title Born Again Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/1805.04770