CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Xinze, Chen, Chen, Yang, Yinfei, Chen, Hong-You, Zhang, Bowen, Pal, Aditya, Zhu, Xiangxin, Du, Xianzhi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918032998137856
author Wang, Xinze
Chen, Chen
Yang, Yinfei
Chen, Hong-You
Zhang, Bowen
Pal, Aditya
Zhu, Xiangxin
Du, Xianzhi
author_facet Wang, Xinze
Chen, Chen
Yang, Yinfei
Chen, Hong-You
Zhang, Bowen
Pal, Aditya
Zhu, Xiangxin
Du, Xianzhi
contents Mixture-of-Experts (MoE) models are crucial for scaling model capacity while controlling inference costs. While integrating MoE into multimodal models like CLIP improves performance, training these models is notoriously challenging and expensive. We propose CLIP-Upcycling (CLIP-UP), an efficient alternative training strategy that converts a pre-trained dense CLIP model into a sparse MoE architecture. Through extensive experimentation with various settings and auxiliary losses, we demonstrate that CLIP-UP significantly reduces training complexity and cost. Remarkably, our sparse CLIP B/16 model, trained with CLIP-UP, outperforms its dense counterpart by 7.2% and 6.6% on COCO and Flickr30k text-to-image Recall@1 benchmarks respectively. It even surpasses the larger CLIP L/14 model on this task while using only 30% of the inference FLOPs. We further demonstrate the generalizability of our training recipe across different scales, establishing sparse upcycling as a practical and scalable approach for building efficient, high-performance CLIP models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00965
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
Wang, Xinze
Chen, Chen
Yang, Yinfei
Chen, Hong-You
Zhang, Bowen
Pal, Aditya
Zhu, Xiangxin
Du, Xianzhi
Computer Vision and Pattern Recognition
Machine Learning
Mixture-of-Experts (MoE) models are crucial for scaling model capacity while controlling inference costs. While integrating MoE into multimodal models like CLIP improves performance, training these models is notoriously challenging and expensive. We propose CLIP-Upcycling (CLIP-UP), an efficient alternative training strategy that converts a pre-trained dense CLIP model into a sparse MoE architecture. Through extensive experimentation with various settings and auxiliary losses, we demonstrate that CLIP-UP significantly reduces training complexity and cost. Remarkably, our sparse CLIP B/16 model, trained with CLIP-UP, outperforms its dense counterpart by 7.2% and 6.6% on COCO and Flickr30k text-to-image Recall@1 benchmarks respectively. It even surpasses the larger CLIP L/14 model on this task while using only 30% of the inference FLOPs. We further demonstrate the generalizability of our training recipe across different scales, establishing sparse upcycling as a practical and scalable approach for building efficient, high-performance CLIP models.
title CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2502.00965