Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lebailly, Tim, Veerabadran, Vijay, Kottur, Satwik, Ridgeway, Karl, Iuzzolino, Michael Louis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912587641257984
author Lebailly, Tim
Veerabadran, Vijay
Kottur, Satwik
Ridgeway, Karl
Iuzzolino, Michael Louis
author_facet Lebailly, Tim
Veerabadran, Vijay
Kottur, Satwik
Ridgeway, Karl
Iuzzolino, Michael Louis
contents Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings indicate. Orthogonal to advancements in generative VLMs, another line of research has focused on representation learning for vision-language alignment, targeting zero-shot inference for dense tasks like segmentation. In this work, we bridge these two directions by densely aligning images with synthetic descriptions generated by VLMs. Synthetic captions are inexpensive, scalable, and easy to generate, making them an excellent source of high-level semantic understanding for dense alignment methods. Empirically, our approach outperforms prior work on standard zero-shot open-vocabulary segmentation benchmarks/datasets, while also being more data-efficient.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
Lebailly, Tim
Veerabadran, Vijay
Kottur, Satwik
Ridgeway, Karl
Iuzzolino, Michael Louis
Computer Vision and Pattern Recognition
Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings indicate. Orthogonal to advancements in generative VLMs, another line of research has focused on representation learning for vision-language alignment, targeting zero-shot inference for dense tasks like segmentation. In this work, we bridge these two directions by densely aligning images with synthetic descriptions generated by VLMs. Synthetic captions are inexpensive, scalable, and easy to generate, making them an excellent source of high-level semantic understanding for dense alignment methods. Empirically, our approach outperforms prior work on standard zero-shot open-vocabulary segmentation benchmarks/datasets, while also being more data-efficient.
title Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.11840