Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jeong, Hyeonsu, Chung, Hye Won
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910835121586176
author Jeong, Hyeonsu
Chung, Hye Won
author_facet Jeong, Hyeonsu
Chung, Hye Won
contents We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical analysis reveals that multi-round self-distillation effectively performs label averaging among instances with high feature correlations, governed by the eigenvectors of the Gram matrix derived from input features. This process leads to clustered predictions and improved generalization, mitigating the impact of label noise by reducing the model's reliance on potentially corrupted labels. We establish conditions under which multi-round self-distillation achieves 100% population accuracy despite label noise. Furthermore, we introduce a novel, efficient single-round self-distillation method using refined partial labels from the teacher's top two softmax outputs, referred to as the PLL student model. This approach replicates the benefits of multi-round distillation in a single round, achieving comparable or superior performance--especially in high-noise scenarios--while significantly reducing computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10482
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels
Jeong, Hyeonsu
Chung, Hye Won
Machine Learning
We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical analysis reveals that multi-round self-distillation effectively performs label averaging among instances with high feature correlations, governed by the eigenvectors of the Gram matrix derived from input features. This process leads to clustered predictions and improved generalization, mitigating the impact of label noise by reducing the model's reliance on potentially corrupted labels. We establish conditions under which multi-round self-distillation achieves 100% population accuracy despite label noise. Furthermore, we introduce a novel, efficient single-round self-distillation method using refined partial labels from the teacher's top two softmax outputs, referred to as the PLL student model. This approach replicates the benefits of multi-round distillation in a single round, achieving comparable or superior performance--especially in high-noise scenarios--while significantly reducing computational cost.
title Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels
topic Machine Learning
url https://arxiv.org/abs/2402.10482