Multimodal Autoregressive Pre-training of Large Vision Encoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fini, Enrico, Shukor, Mustafa, Li, Xiujun, Dufter, Philipp, Klein, Michal, Haldimann, David, Aitharaju, Sai, da Costa, Victor Guilherme Turrisi, Béthune, Louis, Gan, Zhe, Toshev, Alexander T, Eichner, Marcin, Nabi, Moin, Yang, Yinfei, Susskind, Joshua M., El-Nouby, Alaaeldin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910707435438080
author Fini, Enrico
Shukor, Mustafa
Li, Xiujun
Dufter, Philipp
Klein, Michal
Haldimann, David
Aitharaju, Sai
da Costa, Victor Guilherme Turrisi
Béthune, Louis
Gan, Zhe
Toshev, Alexander T
Eichner, Marcin
Nabi, Moin
Yang, Yinfei
Susskind, Joshua M.
El-Nouby, Alaaeldin
author_facet Fini, Enrico
Shukor, Mustafa
Li, Xiujun
Dufter, Philipp
Klein, Michal
Haldimann, David
Aitharaju, Sai
da Costa, Victor Guilherme Turrisi
Béthune, Louis
Gan, Zhe
Toshev, Alexander T
Eichner, Marcin
Nabi, Moin
Yang, Yinfei
Susskind, Joshua M.
El-Nouby, Alaaeldin
contents We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14402
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Autoregressive Pre-training of Large Vision Encoders
Fini, Enrico
Shukor, Mustafa
Li, Xiujun
Dufter, Philipp
Klein, Michal
Haldimann, David
Aitharaju, Sai
da Costa, Victor Guilherme Turrisi
Béthune, Louis
Gan, Zhe
Toshev, Alexander T
Eichner, Marcin
Nabi, Moin
Yang, Yinfei
Susskind, Joshua M.
El-Nouby, Alaaeldin
Computer Vision and Pattern Recognition
Machine Learning
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
title Multimodal Autoregressive Pre-training of Large Vision Encoders
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.14402