MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Truong, Chau, Quang, Hieu Ta, Le, Dung D.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911318187966464
author Truong, Chau
Quang, Hieu Ta
Le, Dung D.
author_facet Truong, Chau
Quang, Hieu Ta
Le, Dung D.
contents Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descriptions. Recent advances mitigate this challenge by leveraging region-proposal information to map visual regions with corresponding sentences from lengthy captions, yet incurring notable deployment costs. We introduce MulCLIP, a novel end-to-end multi-level alignment framework that bridges natural long-text structures with image components. MulCLIP first preserves global contrastive alignment between images and both summary and long captions, while extending positional embeddings for longer text sequences. To further enhance fine-grained understanding, we propose two novel strategies: (1) a token reconstruction alignment over locally calibrated features to strengthen semantic connections between words and image patches, and (2) a subcaption-aggregated patch alignment that automatically extracts and aggregates context-rich patches for each subcaption. Experimental results across diverse benchmarks demonstrate our method consistently improves downstream performance, while ablation studies confirm its multi-scale alignment is the key factor driving better fine-grained capability than region-proposal-assisted approaches, making it particularly suitable for diverse real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07128
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
Truong, Chau
Quang, Hieu Ta
Le, Dung D.
Computer Vision and Pattern Recognition
Vision-language models like CLIP show impressive ability to align images and text, but their training on short, concise captions makes them struggle with lengthy, detailed descriptions. Recent advances mitigate this challenge by leveraging region-proposal information to map visual regions with corresponding sentences from lengthy captions, yet incurring notable deployment costs. We introduce MulCLIP, a novel end-to-end multi-level alignment framework that bridges natural long-text structures with image components. MulCLIP first preserves global contrastive alignment between images and both summary and long captions, while extending positional embeddings for longer text sequences. To further enhance fine-grained understanding, we propose two novel strategies: (1) a token reconstruction alignment over locally calibrated features to strengthen semantic connections between words and image patches, and (2) a subcaption-aggregated patch alignment that automatically extracts and aggregates context-rich patches for each subcaption. Experimental results across diverse benchmarks demonstrate our method consistently improves downstream performance, while ablation studies confirm its multi-scale alignment is the key factor driving better fine-grained capability than region-proposal-assisted approaches, making it particularly suitable for diverse real-world applications.
title MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.07128