PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Campagnolo, Thomas, Malis, Ezio, Martinet, Philippe, Bahl, Gaetan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914070313041920
author Campagnolo, Thomas
Malis, Ezio
Martinet, Philippe
Bahl, Gaetan
author_facet Campagnolo, Thomas
Malis, Ezio
Martinet, Philippe
Bahl, Gaetan
contents Understanding how natural language phrases correspond to specific regions in images is a key challenge in multimodal semantic segmentation. Recent advances in phrase grounding are largely limited to single-view images, neglecting the rich geometric cues available in stereo vision. For this, we introduce PhraseStereo, the first novel dataset that brings phrase-region segmentation to stereo image pairs. PhraseStereo builds upon the PhraseCut dataset by leveraging GenStereo to generate accurate right-view images from existing single-view data, enabling the extension of phrase grounding into the stereo domain. This new setting introduces unique challenges and opportunities for multimodal learning, particularly in leveraging depth cues for more precise and context-aware grounding. By providing stereo image pairs with aligned segmentation masks and phrase annotations, PhraseStereo lays the foundation for future research at the intersection of language, vision, and 3D perception, encouraging the development of models that can reason jointly over semantics and geometry. The PhraseStereo dataset will be released online upon acceptance of this work.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
Campagnolo, Thomas
Malis, Ezio
Martinet, Philippe
Bahl, Gaetan
Computer Vision and Pattern Recognition
Understanding how natural language phrases correspond to specific regions in images is a key challenge in multimodal semantic segmentation. Recent advances in phrase grounding are largely limited to single-view images, neglecting the rich geometric cues available in stereo vision. For this, we introduce PhraseStereo, the first novel dataset that brings phrase-region segmentation to stereo image pairs. PhraseStereo builds upon the PhraseCut dataset by leveraging GenStereo to generate accurate right-view images from existing single-view data, enabling the extension of phrase grounding into the stereo domain. This new setting introduces unique challenges and opportunities for multimodal learning, particularly in leveraging depth cues for more precise and context-aware grounding. By providing stereo image pairs with aligned segmentation masks and phrase annotations, PhraseStereo lays the foundation for future research at the intersection of language, vision, and 3D perception, encouraging the development of models that can reason jointly over semantics and geometry. The PhraseStereo dataset will be released online upon acceptance of this work.
title PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.00818