CoMP: Continual Multimodal Pre-training for Vision Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yitong, Meng, Lingchen, Peng, Wujian, Wu, Zuxuan, Jiang, Yu-Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912380046278656
author Chen, Yitong
Meng, Lingchen
Peng, Wujian
Wu, Zuxuan
Jiang, Yu-Gang
author_facet Chen, Yitong
Meng, Lingchen
Peng, Wujian
Wu, Zuxuan
Jiang, Yu-Gang
contents Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this paper, we continually pre-train prevailing VFMs in a multimodal manner such that they can effortlessly process visual inputs of varying sizes and produce visual representations that are more aligned with language representations, regardless of their original pre-training process. To this end, we introduce CoMP, a carefully designed multimodal pre-training pipeline. CoMP uses a Continual Rotary Position Embedding to accommodate visual inputs with different resolutions, and an Alignment Loss between visual and textual features for better cross-modal alignment. After continual pre-training, leading VFMs like DINOv2, SigLIP and AIMv2 achieve remarkable improvements not only in multimodal understanding tasks but also in generic classification and segmentation tasks. Remarkably, CoMP-AIMv2 achieves scores of 64.9 on ChartQA with a 0.5B LLM, while maintaining an 87.3% accuracy on ImageNet-1K and a 51.8 mIoU on ADE20K under frozen chunk evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoMP: Continual Multimodal Pre-training for Vision Foundation Models
Chen, Yitong
Meng, Lingchen
Peng, Wujian
Wu, Zuxuan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this paper, we continually pre-train prevailing VFMs in a multimodal manner such that they can effortlessly process visual inputs of varying sizes and produce visual representations that are more aligned with language representations, regardless of their original pre-training process. To this end, we introduce CoMP, a carefully designed multimodal pre-training pipeline. CoMP uses a Continual Rotary Position Embedding to accommodate visual inputs with different resolutions, and an Alignment Loss between visual and textual features for better cross-modal alignment. After continual pre-training, leading VFMs like DINOv2, SigLIP and AIMv2 achieve remarkable improvements not only in multimodal understanding tasks but also in generic classification and segmentation tasks. Remarkably, CoMP-AIMv2 achieves scores of 64.9 on ChartQA with a 0.5B LLM, while maintaining an 87.3% accuracy on ImageNet-1K and a 51.8 mIoU on ADE20K under frozen chunk evaluation.
title CoMP: Continual Multimodal Pre-training for Vision Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18931