Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vemulapalli, Raviteja, Pouransari, Hadi, Faghri, Fartash, Mehta, Sachin, Farajtabar, Mehrdad, Rastegari, Mohammad, Tuzel, Oncel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909237250097152
author Vemulapalli, Raviteja
Pouransari, Hadi
Faghri, Fartash
Mehta, Sachin
Farajtabar, Mehrdad
Rastegari, Mohammad
Tuzel, Oncel
author_facet Vemulapalli, Raviteja
Pouransari, Hadi
Faghri, Fartash
Mehta, Sachin
Farajtabar, Mehrdad
Rastegari, Mohammad
Tuzel, Oncel
contents Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot be deployed for many real-world applications. Motivated by this, we ask the following important question, "How can we leverage the knowledge from a large VFM to train a small task-specific model for a new target task with limited labeled training data?", and propose a simple task-oriented knowledge transfer approach as a highly effective solution to this problem. Our experimental results on five target tasks show that the proposed approach outperforms task-agnostic VFM distillation, web-scale CLIP pretraining, supervised ImageNet pretraining, and self-supervised DINO pretraining by up to 11.6%, 22.1%, 13.7%, and 29.8%, respectively. Furthermore, the proposed approach also demonstrates up to 9x, 4x and 15x reduction in pretraining compute cost when compared to task-agnostic VFM distillation, ImageNet pretraining and DINO pretraining, respectively, while outperforming them. We also show that the dataset used for transferring knowledge has a significant effect on the final target task performance, and introduce a retrieval-augmented knowledge transfer strategy that uses web-scale image retrieval to curate effective transfer sets.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18237
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
Vemulapalli, Raviteja
Pouransari, Hadi
Faghri, Fartash
Mehta, Sachin
Farajtabar, Mehrdad
Rastegari, Mohammad
Tuzel, Oncel
Computer Vision and Pattern Recognition
Machine Learning
Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot be deployed for many real-world applications. Motivated by this, we ask the following important question, "How can we leverage the knowledge from a large VFM to train a small task-specific model for a new target task with limited labeled training data?", and propose a simple task-oriented knowledge transfer approach as a highly effective solution to this problem. Our experimental results on five target tasks show that the proposed approach outperforms task-agnostic VFM distillation, web-scale CLIP pretraining, supervised ImageNet pretraining, and self-supervised DINO pretraining by up to 11.6%, 22.1%, 13.7%, and 29.8%, respectively. Furthermore, the proposed approach also demonstrates up to 9x, 4x and 15x reduction in pretraining compute cost when compared to task-agnostic VFM distillation, ImageNet pretraining and DINO pretraining, respectively, while outperforming them. We also show that the dataset used for transferring knowledge has a significant effect on the final target task performance, and introduce a retrieval-augmented knowledge transfer strategy that uses web-scale image retrieval to curate effective transfer sets.
title Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2311.18237