Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cao, Bingyi, Chen, Koert, Maninis, Kevis-Kokitsi, Chen, Kaifeng, Karpur, Arjun, Xia, Ye, Dua, Sahil, Dabral, Tanmaya, Han, Guangxing, Han, Bohyung, Ainslie, Joshua, Bewley, Alex, Jacob, Mithun, Wagner, René, Ramos, Washington, Choromanski, Krzysztof, Seyedhosseini, Mojtaba, Zhou, Howard, Araujo, André
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2604.12012
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911589238571008
author Cao, Bingyi
Chen, Koert
Maninis, Kevis-Kokitsi
Chen, Kaifeng
Karpur, Arjun
Xia, Ye
Dua, Sahil
Dabral, Tanmaya
Han, Guangxing
Han, Bohyung
Ainslie, Joshua
Bewley, Alex
Jacob, Mithun
Wagner, René
Ramos, Washington
Choromanski, Krzysztof
Seyedhosseini, Mojtaba
Zhou, Howard
Araujo, André
author_facet Cao, Bingyi
Chen, Koert
Maninis, Kevis-Kokitsi
Chen, Kaifeng
Karpur, Arjun
Xia, Ye
Dua, Sahil
Dabral, Tanmaya
Han, Guangxing
Han, Bohyung
Ainslie, Joshua
Bewley, Alex
Jacob, Mithun
Wagner, René
Ramos, Washington
Choromanski, Krzysztof
Seyedhosseini, Mojtaba
Zhou, Howard
Araujo, André
contents Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch representations with text embeddings of corresponding concepts. In this work, we investigate this critical issue and propose novel techniques to enhance this capability in foundational vision-language models. First, we reveal that a patch-level distillation procedure significantly boosts dense patch-text alignment -- surprisingly, the patch-text alignment of the distilled student model strongly surpasses that of the teacher model. This observation inspires us to consider modifications to pretraining recipes, leading us to propose iBOT++, an upgrade to the commonly-used iBOT masked image objective, where unmasked tokens also contribute directly to the loss. This dramatically enhances patch-text alignment of pretrained models. Additionally, to improve vision-language pretraining efficiency and effectiveness, we modify the exponential moving average setup in the learning recipe, and introduce a caption sampling strategy to benefit from synthetic captions at different granularities. Combining these components, we develop TIPSv2, a new family of image-text encoder models suitable for a wide range of downstream applications. Through comprehensive experiments on 9 tasks and 20 datasets, we demonstrate strong performance, generally on par with or better than recent vision encoder models. Code and models are released via our project page at https://gdm-tipsv2.github.io/ .
format Preprint
id arxiv_https___arxiv_org_abs_2604_12012
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Cao, Bingyi
Chen, Koert
Maninis, Kevis-Kokitsi
Chen, Kaifeng
Karpur, Arjun
Xia, Ye
Dua, Sahil
Dabral, Tanmaya
Han, Guangxing
Han, Bohyung
Ainslie, Joshua
Bewley, Alex
Jacob, Mithun
Wagner, René
Ramos, Washington
Choromanski, Krzysztof
Seyedhosseini, Mojtaba
Zhou, Howard
Araujo, André
Computer Vision and Pattern Recognition
Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch representations with text embeddings of corresponding concepts. In this work, we investigate this critical issue and propose novel techniques to enhance this capability in foundational vision-language models. First, we reveal that a patch-level distillation procedure significantly boosts dense patch-text alignment -- surprisingly, the patch-text alignment of the distilled student model strongly surpasses that of the teacher model. This observation inspires us to consider modifications to pretraining recipes, leading us to propose iBOT++, an upgrade to the commonly-used iBOT masked image objective, where unmasked tokens also contribute directly to the loss. This dramatically enhances patch-text alignment of pretrained models. Additionally, to improve vision-language pretraining efficiency and effectiveness, we modify the exponential moving average setup in the learning recipe, and introduce a caption sampling strategy to benefit from synthetic captions at different granularities. Combining these components, we develop TIPSv2, a new family of image-text encoder models suitable for a wide range of downstream applications. Through comprehensive experiments on 9 tasks and 20 datasets, we demonstrate strong performance, generally on par with or better than recent vision encoder models. Code and models are released via our project page at https://gdm-tipsv2.github.io/ .
title TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.12012