S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916897798225920 |
|---|---|
| author | Xu, Huihui Ye, Jin Wang, Hongqiu Ji, Changkai Lin, Jiashi Hu, Ming Huang, Ziyan Chen, Ying Ma, Chenglong Li, Tianbin Liu, Lihao He, Junjun Zhu, Lei |
| author_facet | Xu, Huihui Ye, Jin Wang, Hongqiu Ji, Changkai Lin, Jiashi Hu, Ming Huang, Ziyan Chen, Ying Ma, Chenglong Li, Tianbin Liu, Lihao He, Junjun Zhu, Lei |
| contents | Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Our code and pretrained models are available at https://github.com/bio-mlhui/S2-UniSeg |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06995 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision Xu, Huihui Ye, Jin Wang, Hongqiu Ji, Changkai Lin, Jiashi Hu, Ming Huang, Ziyan Chen, Ying Ma, Chenglong Li, Tianbin Liu, Lihao He, Junjun Zhu, Lei Computer Vision and Pattern Recognition Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Our code and pretrained models are available at https://github.com/bio-mlhui/S2-UniSeg |
| title | S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.06995 |