Multimodal Autoregressive Pre-training of Large Vision Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910707435438080 |
|---|---|
| author | Fini, Enrico Shukor, Mustafa Li, Xiujun Dufter, Philipp Klein, Michal Haldimann, David Aitharaju, Sai da Costa, Victor Guilherme Turrisi Béthune, Louis Gan, Zhe Toshev, Alexander T Eichner, Marcin Nabi, Moin Yang, Yinfei Susskind, Joshua M. El-Nouby, Alaaeldin |
| author_facet | Fini, Enrico Shukor, Mustafa Li, Xiujun Dufter, Philipp Klein, Michal Haldimann, David Aitharaju, Sai da Costa, Victor Guilherme Turrisi Béthune, Louis Gan, Zhe Toshev, Alexander T Eichner, Marcin Nabi, Moin Yang, Yinfei Susskind, Joshua M. El-Nouby, Alaaeldin |
| contents | We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_14402 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Multimodal Autoregressive Pre-training of Large Vision Encoders Fini, Enrico Shukor, Mustafa Li, Xiujun Dufter, Philipp Klein, Michal Haldimann, David Aitharaju, Sai da Costa, Victor Guilherme Turrisi Béthune, Louis Gan, Zhe Toshev, Alexander T Eichner, Marcin Nabi, Moin Yang, Yinfei Susskind, Joshua M. El-Nouby, Alaaeldin Computer Vision and Pattern Recognition Machine Learning We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings. |
| title | Multimodal Autoregressive Pre-training of Large Vision Encoders |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2411.14402 |