Saved in:
Bibliographic Details
Main Authors: Popp, Niclas, Metzen, Jan Hendrik, Hein, Matthias
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.16637
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909181872701440
author Popp, Niclas
Metzen, Jan Hendrik
Hein, Matthias
author_facet Popp, Niclas
Metzen, Jan Hendrik
Hein, Matthias
contents Multi-modal foundation models such as CLIP have showcased impressive zero-shot capabilities. However, their applicability in resource-constrained environments is limited due to their large number of parameters and high inference time. While existing approaches have scaled down the entire CLIP architecture, we focus on training smaller variants of the image encoder, which suffices for efficient zero-shot classification. The use of synthetic data has shown promise in distilling representations from larger teachers, resulting in strong few-shot and linear probe performance. However, we find that this approach surprisingly fails in true zero-shot settings when using contrastive losses. We identify the exploitation of spurious features as being responsible for poor generalization between synthetic and real data. However, by using the image feature-based L2 distillation loss, we mitigate these problems and train students that achieve zero-shot performance which on four domain-specific datasets is on-par with a ViT-B/32 teacher model trained on DataCompXL, while featuring up to 92% fewer parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16637
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-Shot Distillation for Image Encoders: How to Make Effective Use of Synthetic Data
Popp, Niclas
Metzen, Jan Hendrik
Hein, Matthias
Computer Vision and Pattern Recognition
Multi-modal foundation models such as CLIP have showcased impressive zero-shot capabilities. However, their applicability in resource-constrained environments is limited due to their large number of parameters and high inference time. While existing approaches have scaled down the entire CLIP architecture, we focus on training smaller variants of the image encoder, which suffices for efficient zero-shot classification. The use of synthetic data has shown promise in distilling representations from larger teachers, resulting in strong few-shot and linear probe performance. However, we find that this approach surprisingly fails in true zero-shot settings when using contrastive losses. We identify the exploitation of spurious features as being responsible for poor generalization between synthetic and real data. However, by using the image feature-based L2 distillation loss, we mitigate these problems and train students that achieve zero-shot performance which on four domain-specific datasets is on-par with a ViT-B/32 teacher model trained on DataCompXL, while featuring up to 92% fewer parameters.
title Zero-Shot Distillation for Image Encoders: How to Make Effective Use of Synthetic Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.16637