CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Baichen, Lyu, Qi, Wang, Xudong, Dong, Jiahua, Liu, Lianqing, Han, Zhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916898512306176
author Liu, Baichen
Lyu, Qi
Wang, Xudong
Dong, Jiahua
Liu, Lianqing
Han, Zhi
author_facet Liu, Baichen
Lyu, Qi
Wang, Xudong
Dong, Jiahua
Liu, Lianqing
Han, Zhi
contents Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal consistency across frames. In this work, we introduce Contrastive Residual Injection and Semantic Prompting (CRISP), an earlier attempt tailored to address the instance-wise, category-wise, and task-wise confusion in continual video instance segmentation. For instance-wise learning, we model instance tracking and construct instance correlation loss, which emphasizes the correlation with the prior query space while strengthening the specificity of the current task query. For category-wise learning, we build an adaptive residual semantic prompt (ARSP) learning framework, which constructs a learnable semantic residual prompt pool generated by category text and uses an adjustive query-prompt matching mechanism to build a mapping relationship between the query of the current task and the semantic residual prompt. Meanwhile, a semantic consistency loss based on the contrastive learning is introduced to maintain semantic coherence between object queries and residual prompts during incremental training. For task-wise learning, to ensure the correlation at the inter-task level within the query space, we introduce a concise yet powerful initialization strategy for incremental prompts. Extensive experiments on YouTube-VIS-2019 and YouTube-VIS-2021 datasets demonstrate that CRISP significantly outperforms existing continual segmentation methods in the long-term continual video instance segmentation task, avoiding catastrophic forgetting and effectively improving segmentation and classification performance. The code is available at https://github.com/01upup10/CRISP.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
Liu, Baichen
Lyu, Qi
Wang, Xudong
Dong, Jiahua
Liu, Lianqing
Han, Zhi
Computer Vision and Pattern Recognition
Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal consistency across frames. In this work, we introduce Contrastive Residual Injection and Semantic Prompting (CRISP), an earlier attempt tailored to address the instance-wise, category-wise, and task-wise confusion in continual video instance segmentation. For instance-wise learning, we model instance tracking and construct instance correlation loss, which emphasizes the correlation with the prior query space while strengthening the specificity of the current task query. For category-wise learning, we build an adaptive residual semantic prompt (ARSP) learning framework, which constructs a learnable semantic residual prompt pool generated by category text and uses an adjustive query-prompt matching mechanism to build a mapping relationship between the query of the current task and the semantic residual prompt. Meanwhile, a semantic consistency loss based on the contrastive learning is introduced to maintain semantic coherence between object queries and residual prompts during incremental training. For task-wise learning, to ensure the correlation at the inter-task level within the query space, we introduce a concise yet powerful initialization strategy for incremental prompts. Extensive experiments on YouTube-VIS-2019 and YouTube-VIS-2021 datasets demonstrate that CRISP significantly outperforms existing continual segmentation methods in the long-term continual video instance segmentation task, avoiding catastrophic forgetting and effectively improving segmentation and classification performance. The code is available at https://github.com/01upup10/CRISP.
title CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.10432