Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jin, Zhang, Bingfeng, Pang, Jian, Liu, Mengyu, Chen, Honglong, Liu, Weifeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914165640134656
author Wang, Jin
Zhang, Bingfeng
Pang, Jian
Liu, Mengyu
Chen, Honglong
Liu, Weifeng
author_facet Wang, Jin
Zhang, Bingfeng
Pang, Jian
Liu, Mengyu
Chen, Honglong
Liu, Weifeng
contents Few-shot segmentation (FSS) aims to segment novel classes under the guidance of limited support samples by a meta-learning paradigm. Existing methods mainly mine references from support images as meta guidance. However, due to intra-class variations among visual representations, the meta information extracted from support images cannot produce accurate guidance to segment untrained classes. In this paper, we argue that the references from support images may not be essential, the key to the support role is to provide unbiased meta guidance for both trained and untrained classes. We then introduce a Language-Driven Attribute Generalization (LDAG) architecture to utilize inherent target property language descriptions to build robust support strategy. Specifically, to obtain an unbiased support representation, we design a Multi-attribute Enhancement (MaE) module, which produces multiple detailed attribute descriptions of the target class through Large Language Models (LLMs), and then builds refined visual-text prior guidance utilizing multi-modal matching. Meanwhile, due to text-vision modal shift, attribute text struggles to promote visual feature representation, we design a Multi-modal Attribute Alignment (MaA) to achieve cross-modal interaction between attribute texts and visual feature. Experiments show that our proposed method outperforms existing approaches by a clear margin and achieves the new state-of-the art performance. The code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16435
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
Wang, Jin
Zhang, Bingfeng
Pang, Jian
Liu, Mengyu
Chen, Honglong
Liu, Weifeng
Computer Vision and Pattern Recognition
Few-shot segmentation (FSS) aims to segment novel classes under the guidance of limited support samples by a meta-learning paradigm. Existing methods mainly mine references from support images as meta guidance. However, due to intra-class variations among visual representations, the meta information extracted from support images cannot produce accurate guidance to segment untrained classes. In this paper, we argue that the references from support images may not be essential, the key to the support role is to provide unbiased meta guidance for both trained and untrained classes. We then introduce a Language-Driven Attribute Generalization (LDAG) architecture to utilize inherent target property language descriptions to build robust support strategy. Specifically, to obtain an unbiased support representation, we design a Multi-attribute Enhancement (MaE) module, which produces multiple detailed attribute descriptions of the target class through Large Language Models (LLMs), and then builds refined visual-text prior guidance utilizing multi-modal matching. Meanwhile, due to text-vision modal shift, attribute text struggles to promote visual feature representation, we design a Multi-modal Attribute Alignment (MaA) to achieve cross-modal interaction between attribute texts and visual feature. Experiments show that our proposed method outperforms existing approaches by a clear margin and achieves the new state-of-the art performance. The code will be released.
title Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16435