Saved in:
Bibliographic Details
Main Authors: Marrie, Juliette, Arbel, Michael, Mairal, Julien, Larlus, Diane
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.11305
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909193321054208
author Marrie, Juliette
Arbel, Michael
Mairal, Julien
Larlus, Diane
author_facet Marrie, Juliette
Arbel, Michael
Mairal, Julien
Larlus, Diane
contents Large pretrained visual models exhibit remarkable generalization across diverse recognition tasks. Yet, real-world applications often demand compact models tailored to specific problems. Variants of knowledge distillation have been devised for such a purpose, enabling task-specific compact models (the students) to learn from a generic large pretrained one (the teacher). In this paper, we show that the excellent robustness and versatility of recent pretrained models challenge common practices established in the literature, calling for a new set of optimal guidelines for task-specific distillation. To address the lack of samples in downstream tasks, we also show that a variant of Mixup based on stable diffusion complements standard data augmentation. This strategy eliminates the need for engineered text prompts and improves distillation of generic models into streamlined specialized networks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11305
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Good Practices for Task-Specific Distillation of Large Pretrained Visual Models
Marrie, Juliette
Arbel, Michael
Mairal, Julien
Larlus, Diane
Computer Vision and Pattern Recognition
Large pretrained visual models exhibit remarkable generalization across diverse recognition tasks. Yet, real-world applications often demand compact models tailored to specific problems. Variants of knowledge distillation have been devised for such a purpose, enabling task-specific compact models (the students) to learn from a generic large pretrained one (the teacher). In this paper, we show that the excellent robustness and versatility of recent pretrained models challenge common practices established in the literature, calling for a new set of optimal guidelines for task-specific distillation. To address the lack of samples in downstream tasks, we also show that a variant of Mixup based on stable diffusion complements standard data augmentation. This strategy eliminates the need for engineered text prompts and improves distillation of generic models into streamlined specialized networks.
title On Good Practices for Task-Specific Distillation of Large Pretrained Visual Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.11305