Data-efficient Large Vision Models through Sequential Autoregression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Jianyuan, Hao, Zhiwei, Wang, Chengcheng, Tang, Yehui, Wu, Han, Hu, Han, Han, Kai, Xu, Chang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910473768665088
author Guo, Jianyuan
Hao, Zhiwei
Wang, Chengcheng
Tang, Yehui
Wu, Han
Hu, Han
Han, Kai
Xu, Chang
author_facet Guo, Jianyuan
Hao, Zhiwei
Wang, Chengcheng
Tang, Yehui
Wu, Han
Hu, Han
Han, Kai
Xu, Chang
contents Training general-purpose vision models on purely sequential visual data, eschewing linguistic inputs, has heralded a new frontier in visual understanding. These models are intended to not only comprehend but also seamlessly transit to out-of-domain tasks. However, current endeavors are hamstrung by an over-reliance on colossal models, exemplified by models with upwards of 3B parameters, and the necessity for an extensive corpus of visual data, often comprising a staggering 400B tokens. In this paper, we delve into the development of an efficient, autoregression-based vision model, innovatively architected to operate on a limited dataset. We meticulously demonstrate how this model achieves proficiency in a spectrum of visual tasks spanning both high-level and low-level semantic understanding during the testing phase. Our empirical evaluations underscore the model's agility in adapting to various tasks, heralding a significant reduction in the parameter footprint, and a marked decrease in training data requirements, thereby paving the way for more sustainable and accessible advancements in the field of generalist vision models. The code is available at https://github.com/ggjy/DeLVM.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04841
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data-efficient Large Vision Models through Sequential Autoregression
Guo, Jianyuan
Hao, Zhiwei
Wang, Chengcheng
Tang, Yehui
Wu, Han
Hu, Han
Han, Kai
Xu, Chang
Computer Vision and Pattern Recognition
Training general-purpose vision models on purely sequential visual data, eschewing linguistic inputs, has heralded a new frontier in visual understanding. These models are intended to not only comprehend but also seamlessly transit to out-of-domain tasks. However, current endeavors are hamstrung by an over-reliance on colossal models, exemplified by models with upwards of 3B parameters, and the necessity for an extensive corpus of visual data, often comprising a staggering 400B tokens. In this paper, we delve into the development of an efficient, autoregression-based vision model, innovatively architected to operate on a limited dataset. We meticulously demonstrate how this model achieves proficiency in a spectrum of visual tasks spanning both high-level and low-level semantic understanding during the testing phase. Our empirical evaluations underscore the model's agility in adapting to various tasks, heralding a significant reduction in the parameter footprint, and a marked decrease in training data requirements, thereby paving the way for more sustainable and accessible advancements in the field of generalist vision models. The code is available at https://github.com/ggjy/DeLVM.
title Data-efficient Large Vision Models through Sequential Autoregression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.04841