Unified Open-World Segmentation with Multi-Modal Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yang, Yin, Yufei, Jing, Chenchen, Zhu, Muzhi, Chen, Hao, Xi, Yuling, Feng, Bo, Wang, Hao, Li, Shiyu, Shen, Chunhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908588917653504
author Liu, Yang
Yin, Yufei
Jing, Chenchen
Zhu, Muzhi
Chen, Hao
Xi, Yuling
Feng, Bo
Wang, Hao
Li, Shiyu
Shen, Chunhua
author_facet Liu, Yang
Yin, Yufei
Jing, Chenchen
Zhu, Muzhi
Chen, Hao
Xi, Yuling
Feng, Bo
Wang, Hao
Li, Shiyu
Shen, Chunhua
contents In this work, we present COSINE, a unified open-world segmentation model that consolidates open-vocabulary segmentation and in-context segmentation with multi-modal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi-modal prompts, and a SegDecoder to align these representations, model their interaction, and obtain masks specified by input prompts across different granularities. In this way, COSINE overcomes architectural discrepancies, divergent learning objectives, and distinct representation learning strategies of previous pipelines for open-vocabulary segmentation and in-context segmentation. Comprehensive experiments demonstrate that COSINE has significant performance improvements in both open-vocabulary and in-context segmentation tasks. Our exploratory analyses highlight that the synergistic collaboration between using visual and textual prompts leads to significantly improved generalization over single-modality approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10524
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Open-World Segmentation with Multi-Modal Prompts
Liu, Yang
Yin, Yufei
Jing, Chenchen
Zhu, Muzhi
Chen, Hao
Xi, Yuling
Feng, Bo
Wang, Hao
Li, Shiyu
Shen, Chunhua
Computer Vision and Pattern Recognition
In this work, we present COSINE, a unified open-world segmentation model that consolidates open-vocabulary segmentation and in-context segmentation with multi-modal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi-modal prompts, and a SegDecoder to align these representations, model their interaction, and obtain masks specified by input prompts across different granularities. In this way, COSINE overcomes architectural discrepancies, divergent learning objectives, and distinct representation learning strategies of previous pipelines for open-vocabulary segmentation and in-context segmentation. Comprehensive experiments demonstrate that COSINE has significant performance improvements in both open-vocabulary and in-context segmentation tasks. Our exploratory analyses highlight that the synergistic collaboration between using visual and textual prompts leads to significantly improved generalization over single-modality approaches.
title Unified Open-World Segmentation with Multi-Modal Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10524