Saved in:
Bibliographic Details
Main Authors: Seutin, Corentin, Ettaki, Mohamed Amine, Clément, Michaël, Coupé, Pierrick, Giraud, Rémi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.28348
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914608419176448
author Seutin, Corentin
Ettaki, Mohamed Amine
Clément, Michaël
Coupé, Pierrick
Giraud, Rémi
author_facet Seutin, Corentin
Ettaki, Mohamed Amine
Clément, Michaël
Coupé, Pierrick
Giraud, Rémi
contents Vision-language segmentation models have recently achieved strong performance by leveraging high-level semantic object categories expressed in natural language. However, this semantic dependence limits their ability to reason about intrinsic visual properties such as shape, geometry, or texture, which are essential in many real-world applications. In this work, we introduce Semantic-Agnostic aNd Shape-Aware (SANSA) segmentation, a new paradigm that requires segmentation models to operate solely from non-semantic textual descriptions. To this end, we propose two strategies to generate SANSA segmentation prompts based on either dictionary constraints or example guidance, both generating semantic-agnostic textual descriptions. These prompts are then used to finetune segmentation models under semantic-agnostic supervision. Experiments show that finetuning on SANSA prompts yields up to a 20% mIoU improvement on this new segmentation task, compared to pretrained state-of-the-art models, while maintaining strong performance on standard semantic prompts. These results highlight the importance of low- and mid-level visual reasoning for improving the generalization and controllability of vision-language segmentation models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28348
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
Seutin, Corentin
Ettaki, Mohamed Amine
Clément, Michaël
Coupé, Pierrick
Giraud, Rémi
Computer Vision and Pattern Recognition
Vision-language segmentation models have recently achieved strong performance by leveraging high-level semantic object categories expressed in natural language. However, this semantic dependence limits their ability to reason about intrinsic visual properties such as shape, geometry, or texture, which are essential in many real-world applications. In this work, we introduce Semantic-Agnostic aNd Shape-Aware (SANSA) segmentation, a new paradigm that requires segmentation models to operate solely from non-semantic textual descriptions. To this end, we propose two strategies to generate SANSA segmentation prompts based on either dictionary constraints or example guidance, both generating semantic-agnostic textual descriptions. These prompts are then used to finetune segmentation models under semantic-agnostic supervision. Experiments show that finetuning on SANSA prompts yields up to a 20% mIoU improvement on this new segmentation task, compared to pretrained state-of-the-art models, while maintaining strong performance on standard semantic prompts. These results highlight the importance of low- and mid-level visual reasoning for improving the generalization and controllability of vision-language segmentation models.
title Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.28348