Classification Done Right for Vision-Language Pre-Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Zilong, Ye, Qinghao, Kang, Bingyi, Feng, Jiashi, Fan, Haoqi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909378877063168
author Huang, Zilong
Ye, Qinghao
Kang, Bingyi
Feng, Jiashi
Fan, Haoqi
author_facet Huang, Zilong
Ye, Qinghao
Kang, Bingyi
Feng, Jiashi
Fan, Haoqi
contents We introduce SuperClass, a super simple classification method for vision-language pre-training on image-text data. Unlike its contrastive counterpart CLIP who contrast with a text encoder, SuperClass directly utilizes tokenized raw text as supervised classification labels, without the need for additional text filtering or selection. Due to the absence of the text encoding as contrastive target, SuperClass does not require a text encoder and does not need to maintain a large batch size as CLIP does. SuperClass demonstrated superior performance on various downstream tasks, including classic computer vision benchmarks and vision language downstream tasks. We further explored the scaling behavior of SuperClass on model size, training length, or data size, and reported encouraging results and comparisons to CLIP. https://github.com/x-cls/superclass
format Preprint
id arxiv_https___arxiv_org_abs_2411_03313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Classification Done Right for Vision-Language Pre-Training
Huang, Zilong
Ye, Qinghao
Kang, Bingyi
Feng, Jiashi
Fan, Haoqi
Computer Vision and Pattern Recognition
We introduce SuperClass, a super simple classification method for vision-language pre-training on image-text data. Unlike its contrastive counterpart CLIP who contrast with a text encoder, SuperClass directly utilizes tokenized raw text as supervised classification labels, without the need for additional text filtering or selection. Due to the absence of the text encoding as contrastive target, SuperClass does not require a text encoder and does not need to maintain a large batch size as CLIP does. SuperClass demonstrated superior performance on various downstream tasks, including classic computer vision benchmarks and vision language downstream tasks. We further explored the scaling behavior of SuperClass on model size, training length, or data size, and reported encouraging results and comparisons to CLIP. https://github.com/x-cls/superclass
title Classification Done Right for Vision-Language Pre-Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.03313