VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909245084008448 |
|---|---|
| author | Chu, Xiangxiang Su, Jianlin Zhang, Bo Shen, Chunhua |
| author_facet | Chu, Xiangxiang Su, Jianlin Zhang, Bo Shen, Chunhua |
| contents | Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_00522 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks Chu, Xiangxiang Su, Jianlin Zhang, Bo Shen, Chunhua Computer Vision and Pattern Recognition Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA. |
| title | VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2403.00522 |