See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916022699687936 |
|---|---|
| author | Sun, Boyuan Yin, Bowen Li, Yuanming Wei, Xihan Hou, Qibin |
| author_facet | Sun, Boyuan Yin, Bowen Li, Yuanming Wei, Xihan Hou, Qibin |
| contents | We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual prompts, such as masks or points, SWIM leverages mask supervision only during training to guide cross-modal attention, allowing the model to automatically attend to the user-specified object at inference. Our cross-attention analysis of pretrained multimodal large languagemodels (MLLMs) reveals a systematic discrepancy: Attribute words produce sharp, localized activations in the visual modality, whereas object nouns yield diffuse and scattered patterns due to semantic reference bias and distributed high-level representations. To address this misalignment, we construct NL-Refer, an enriched dataset, in which each object mask is paired with a precise natural language referring expression. SWIM extracts multi-layer cross-attention maps from object nouns and enforces spatial consistency with ground-truth masks. Experimental results demonstrate that SWIM substantially improves text-visual alignment and achieves superior performance over visual-prompt-based methods on fine-grained object understanding benchmarks. The code and data are available at \href{https://github.com/HumanMLLM/SWIM}{https://github.com/HumanMLLM/SWIM}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_18018 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Sun, Boyuan Yin, Bowen Li, Yuanming Wei, Xihan Hou, Qibin Computer Vision and Pattern Recognition Artificial Intelligence Human-Computer Interaction We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual prompts, such as masks or points, SWIM leverages mask supervision only during training to guide cross-modal attention, allowing the model to automatically attend to the user-specified object at inference. Our cross-attention analysis of pretrained multimodal large languagemodels (MLLMs) reveals a systematic discrepancy: Attribute words produce sharp, localized activations in the visual modality, whereas object nouns yield diffuse and scattered patterns due to semantic reference bias and distributed high-level representations. To address this misalignment, we construct NL-Refer, an enriched dataset, in which each object mask is paired with a precise natural language referring expression. SWIM extracts multi-layer cross-attention maps from object nouns and enforces spatial consistency with ground-truth masks. Experimental results demonstrate that SWIM substantially improves text-visual alignment and achieves superior performance over visual-prompt-based methods on fine-grained object understanding benchmarks. The code and data are available at \href{https://github.com/HumanMLLM/SWIM}{https://github.com/HumanMLLM/SWIM}. |
| title | See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Human-Computer Interaction |
| url | https://arxiv.org/abs/2605.18018 |