VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chu, Xiangxiang, Su, Jianlin, Zhang, Bo, Shen, Chunhua
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909245084008448
author Chu, Xiangxiang
Su, Jianlin
Zhang, Bo
Shen, Chunhua
author_facet Chu, Xiangxiang
Su, Jianlin
Zhang, Bo
Shen, Chunhua
contents Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00522
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
Chu, Xiangxiang
Su, Jianlin
Zhang, Bo
Shen, Chunhua
Computer Vision and Pattern Recognition
Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA.
title VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.00522