S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Huihui, Ye, Jin, Wang, Hongqiu, Ji, Changkai, Lin, Jiashi, Hu, Ming, Huang, Ziyan, Chen, Ying, Ma, Chenglong, Li, Tianbin, Liu, Lihao, He, Junjun, Zhu, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916897798225920
author Xu, Huihui
Ye, Jin
Wang, Hongqiu
Ji, Changkai
Lin, Jiashi
Hu, Ming
Huang, Ziyan
Chen, Ying
Ma, Chenglong
Li, Tianbin
Liu, Lihao
He, Junjun
Zhu, Lei
author_facet Xu, Huihui
Ye, Jin
Wang, Hongqiu
Ji, Changkai
Lin, Jiashi
Hu, Ming
Huang, Ziyan
Chen, Ying
Ma, Chenglong
Li, Tianbin
Liu, Lihao
He, Junjun
Zhu, Lei
contents Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Our code and pretrained models are available at https://github.com/bio-mlhui/S2-UniSeg
format Preprint
id arxiv_https___arxiv_org_abs_2508_06995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision
Xu, Huihui
Ye, Jin
Wang, Hongqiu
Ji, Changkai
Lin, Jiashi
Hu, Ming
Huang, Ziyan
Chen, Ying
Ma, Chenglong
Li, Tianbin
Liu, Lihao
He, Junjun
Zhu, Lei
Computer Vision and Pattern Recognition
Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Our code and pretrained models are available at https://github.com/bio-mlhui/S2-UniSeg
title S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06995