Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zichao, Xie, Cihang, Cubuk, Ekin Dogus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914756280975360
author Li, Zichao
Xie, Cihang
Cubuk, Ekin Dogus
author_facet Li, Zichao
Xie, Cihang
Cubuk, Ekin Dogus
contents This paper investigates the performance of the Contrastive Language-Image Pre-training (CLIP) when scaled down to limited computation budgets. We explore CLIP along three dimensions: data, architecture, and training strategies. With regards to data, we demonstrate the significance of high-quality training data and show that a smaller dataset of high-quality data can outperform a larger dataset with lower quality. We also examine how model performance varies with different dataset sizes, suggesting that smaller ViT models are better suited for smaller datasets, while larger models perform better on larger datasets with fixed compute. Additionally, we provide guidance on when to choose a CNN-based architecture or a ViT-based architecture for CLIP training. We compare four CLIP training strategies - SLIP, FLIP, CLIP, and CLIP+Data Augmentation - and show that the choice of training strategy depends on the available compute resource. Our analysis reveals that CLIP+Data Augmentation can achieve comparable performance to CLIP using only half of the training data. This work provides practical insights into how to effectively train and deploy CLIP models, making them more accessible and affordable for practical use in various applications.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08197
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
Li, Zichao
Xie, Cihang
Cubuk, Ekin Dogus
Computer Vision and Pattern Recognition
This paper investigates the performance of the Contrastive Language-Image Pre-training (CLIP) when scaled down to limited computation budgets. We explore CLIP along three dimensions: data, architecture, and training strategies. With regards to data, we demonstrate the significance of high-quality training data and show that a smaller dataset of high-quality data can outperform a larger dataset with lower quality. We also examine how model performance varies with different dataset sizes, suggesting that smaller ViT models are better suited for smaller datasets, while larger models perform better on larger datasets with fixed compute. Additionally, we provide guidance on when to choose a CNN-based architecture or a ViT-based architecture for CLIP training. We compare four CLIP training strategies - SLIP, FLIP, CLIP, and CLIP+Data Augmentation - and show that the choice of training strategy depends on the available compute resource. Our analysis reveals that CLIP+Data Augmentation can achieve comparable performance to CLIP using only half of the training data. This work provides practical insights into how to effectively train and deploy CLIP models, making them more accessible and affordable for practical use in various applications.
title Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.08197