Omni-Referring Image Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zheng, Qiancheng, Shen, Yunhang, Luo, Gen, Song, Baiyang, Sun, Xing, Sun, Xiaoshuai, Zhou, Yiyi, Ji, Rongrong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912753419026432
author Zheng, Qiancheng
Shen, Yunhang
Luo, Gen
Song, Baiyang
Sun, Xing
Sun, Xiaoshuai
Zhou, Yiyi
Ji, Rongrong
author_facet Zheng, Qiancheng
Shen, Yunhang
Luo, Gen
Song, Baiyang
Sun, Xing
Sun, Xiaoshuai
Zhou, Yiyi
Ji, Rongrong
contents In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally conditioned segmentation tasks, such as RIS and visual RIS, OmniRIS supports the input of text instructions and reference images with masks, boxes or scribbles as omni-prompts. This property makes it can well exploit the intrinsic merits of both text and visual modalities, i.e., granular attribute referring and uncommon object grounding, respectively. Besides, OmniRIS can also handle various segmentation settings, such as one v.s. many and many v.s. many, further facilitating its practical use. To promote the research of OmniRIS, we also rigorously design and construct a large dataset termed OmniRef, which consists of 186,939 omni-prompts for 30,956 images, and establish a comprehensive evaluation system. Moreover, a strong and general baseline termed OmniSegNet is also proposed to tackle the key challenges of OmniRIS, such as omni-prompt encoding. The extensive experiments not only validate the capability of OmniSegNet in following omni-modal instructions, but also show the superiority of OmniRIS for highly generalized image segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Omni-Referring Image Segmentation
Zheng, Qiancheng
Shen, Yunhang
Luo, Gen
Song, Baiyang
Sun, Xing
Sun, Xiaoshuai
Zhou, Yiyi
Ji, Rongrong
Computer Vision and Pattern Recognition
In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally conditioned segmentation tasks, such as RIS and visual RIS, OmniRIS supports the input of text instructions and reference images with masks, boxes or scribbles as omni-prompts. This property makes it can well exploit the intrinsic merits of both text and visual modalities, i.e., granular attribute referring and uncommon object grounding, respectively. Besides, OmniRIS can also handle various segmentation settings, such as one v.s. many and many v.s. many, further facilitating its practical use. To promote the research of OmniRIS, we also rigorously design and construct a large dataset termed OmniRef, which consists of 186,939 omni-prompts for 30,956 images, and establish a comprehensive evaluation system. Moreover, a strong and general baseline termed OmniSegNet is also proposed to tackle the key challenges of OmniRIS, such as omni-prompt encoding. The extensive experiments not only validate the capability of OmniSegNet in following omni-modal instructions, but also show the superiority of OmniRIS for highly generalized image segmentation.
title Omni-Referring Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06862