Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salazar, Israfel, Elliott, Desmond, Kementchedjhieva, Yova
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910210094792704
author Salazar, Israfel
Elliott, Desmond
Kementchedjhieva, Yova
author_facet Salazar, Israfel
Elliott, Desmond
Kementchedjhieva, Yova
contents Contrastive vision-language models (VLMs) have made significant progress in binding visual and textual information, yet understanding long, compositional captions remains an open challenge. While these capabilities are often assumed to be closely related, the conditions under which they reinforce each other remain unclear. In this paper, we empirically analyze when compositional reasoning and long-caption understanding transfer across tasks, and when this relationship fails. Through controlled experiments across diverse training objectives, datasets, and architectural designs, we find a bidirectional but sensitive relationship between the two capabilities. Models trained on poorly grounded captions or with limited parameter updates fail to generalize, while high-quality long-caption data with strong visual grounding promotes both capabilities simultaneously. We further show that architectural choices aimed at preserving general alignment, such as frozen positional embeddings, can inadvertently limit compositional learning. Our analysis provides actionable guidelines for data selection and model design to improve VLM generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
Salazar, Israfel
Elliott, Desmond
Kementchedjhieva, Yova
Computer Vision and Pattern Recognition
Contrastive vision-language models (VLMs) have made significant progress in binding visual and textual information, yet understanding long, compositional captions remains an open challenge. While these capabilities are often assumed to be closely related, the conditions under which they reinforce each other remain unclear. In this paper, we empirically analyze when compositional reasoning and long-caption understanding transfer across tasks, and when this relationship fails. Through controlled experiments across diverse training objectives, datasets, and architectural designs, we find a bidirectional but sensitive relationship between the two capabilities. Models trained on poorly grounded captions or with limited parameter updates fail to generalize, while high-quality long-caption data with strong visual grounding promotes both capabilities simultaneously. We further show that architectural choices aimed at preserving general alignment, such as frozen positional embeddings, can inadvertently limit compositional learning. Our analysis provides actionable guidelines for data selection and model design to improve VLM generalization.
title Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.19207