Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.11305 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909193321054208 |
|---|---|
| author | Marrie, Juliette Arbel, Michael Mairal, Julien Larlus, Diane |
| author_facet | Marrie, Juliette Arbel, Michael Mairal, Julien Larlus, Diane |
| contents | Large pretrained visual models exhibit remarkable generalization across diverse recognition tasks. Yet, real-world applications often demand compact models tailored to specific problems. Variants of knowledge distillation have been devised for such a purpose, enabling task-specific compact models (the students) to learn from a generic large pretrained one (the teacher). In this paper, we show that the excellent robustness and versatility of recent pretrained models challenge common practices established in the literature, calling for a new set of optimal guidelines for task-specific distillation. To address the lack of samples in downstream tasks, we also show that a variant of Mixup based on stable diffusion complements standard data augmentation. This strategy eliminates the need for engineered text prompts and improves distillation of generic models into streamlined specialized networks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_11305 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | On Good Practices for Task-Specific Distillation of Large Pretrained Visual Models Marrie, Juliette Arbel, Michael Mairal, Julien Larlus, Diane Computer Vision and Pattern Recognition Large pretrained visual models exhibit remarkable generalization across diverse recognition tasks. Yet, real-world applications often demand compact models tailored to specific problems. Variants of knowledge distillation have been devised for such a purpose, enabling task-specific compact models (the students) to learn from a generic large pretrained one (the teacher). In this paper, we show that the excellent robustness and versatility of recent pretrained models challenge common practices established in the literature, calling for a new set of optimal guidelines for task-specific distillation. To address the lack of samples in downstream tasks, we also show that a variant of Mixup based on stable diffusion complements standard data augmentation. This strategy eliminates the need for engineered text prompts and improves distillation of generic models into streamlined specialized networks. |
| title | On Good Practices for Task-Specific Distillation of Large Pretrained Visual Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2402.11305 |