Rethinking Query-based Transformer for Continual Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yuchen, Shi, Cheng, Wang, Dingyou, Tang, Jiajin, Wei, Zhengxuan, Wu, Yu, Li, Guanbin, Yang, Sibei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915382169698304
author Zhu, Yuchen
Shi, Cheng
Wang, Dingyou
Tang, Jiajin
Wei, Zhengxuan
Wu, Yu
Li, Guanbin
Yang, Sibei
author_facet Zhu, Yuchen
Shi, Cheng
Wang, Dingyou
Tang, Jiajin
Wei, Zhengxuan
Wu, Yu
Li, Guanbin
Yang, Sibei
contents Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods often decouple mask generation from the continual learning process. This study, however, identifies two key issues with decoupled frameworks: loss of plasticity and heavy reliance on input data order. To address these, we conduct an in-depth investigation of the built-in objectness and find that highly aggregated image features provide a shortcut for queries to generate masks through simple feature alignment. Based on this, we propose SimCIS, a simple yet powerful baseline for CIS. Its core idea is to directly select image features for query assignment, ensuring "perfect alignment" to preserve objectness, while simultaneously allowing queries to select new classes to promote plasticity. To further combat catastrophic forgetting of categories, we introduce cross-stage consistency in selection and an innovative "visual query"-based replay mechanism. Experiments demonstrate that SimCIS consistently outperforms state-of-the-art methods across various segmentation tasks, settings, splits, and input data orders. All models and codes will be made publicly available at https://github.com/SooLab/SimCIS.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Query-based Transformer for Continual Image Segmentation
Zhu, Yuchen
Shi, Cheng
Wang, Dingyou
Tang, Jiajin
Wei, Zhengxuan
Wu, Yu
Li, Guanbin
Yang, Sibei
Computer Vision and Pattern Recognition
Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods often decouple mask generation from the continual learning process. This study, however, identifies two key issues with decoupled frameworks: loss of plasticity and heavy reliance on input data order. To address these, we conduct an in-depth investigation of the built-in objectness and find that highly aggregated image features provide a shortcut for queries to generate masks through simple feature alignment. Based on this, we propose SimCIS, a simple yet powerful baseline for CIS. Its core idea is to directly select image features for query assignment, ensuring "perfect alignment" to preserve objectness, while simultaneously allowing queries to select new classes to promote plasticity. To further combat catastrophic forgetting of categories, we introduce cross-stage consistency in selection and an innovative "visual query"-based replay mechanism. Experiments demonstrate that SimCIS consistently outperforms state-of-the-art methods across various segmentation tasks, settings, splits, and input data orders. All models and codes will be made publicly available at https://github.com/SooLab/SimCIS.
title Rethinking Query-based Transformer for Continual Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.07831