Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chenhao, Ji, Yingrui, Meng, Yu, Zhu, Yao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909046969204736
author Wang, Chenhao
Ji, Yingrui
Meng, Yu
Zhu, Yao
author_facet Wang, Chenhao
Ji, Yingrui
Meng, Yu
Zhu, Yao
contents Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded as holistic sentences that entangle multiple semantic units. We propose a Decomposed Vision-Language Alignment framework that explicitly factorizes textual prompts into a concept token and multiple attribute tokens, enabling separate cross-modal interactions for each semantic unit. At the feature level, we introduce a Feature-Gated Cross-Attention module that generates attribute-specific gating maps to fuse information in a multiplicative manner, effectively enforcing compositional semantics. At the scoring level, per-token similarities are aggregated in log-space, producing a stable and interpretable compositional matching. The method can be seamlessly integrated into existing transformer-based segmentation architectures and significantly improves generalization to unseen attribute-category compositions in fine-grained open-vocabulary segmentation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15942
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
Wang, Chenhao
Ji, Yingrui
Meng, Yu
Zhu, Yao
Computer Vision and Pattern Recognition
Artificial Intelligence
Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded as holistic sentences that entangle multiple semantic units. We propose a Decomposed Vision-Language Alignment framework that explicitly factorizes textual prompts into a concept token and multiple attribute tokens, enabling separate cross-modal interactions for each semantic unit. At the feature level, we introduce a Feature-Gated Cross-Attention module that generates attribute-specific gating maps to fuse information in a multiplicative manner, effectively enforcing compositional semantics. At the scoring level, per-token similarities are aggregated in log-space, producing a stable and interpretable compositional matching. The method can be seamlessly integrated into existing transformer-based segmentation architectures and significantly improves generalization to unseen attribute-category compositions in fine-grained open-vocabulary segmentation benchmarks.
title Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.15942