Revisiting Continual Semantic Segmentation with Pre-trained Vision Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Duzhen, Ren, Yong, Cong, Wei, Zheng, Junhao, Su, Qiaoyi, Jia, Shuncheng, Li, Zhong-Zhi, Zhao, Xuanle, Bai, Ye, Chen, Feilong, Tian, Qi, Zhang, Tielin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912523070996480
author Zhang, Duzhen
Ren, Yong
Cong, Wei
Zheng, Junhao
Su, Qiaoyi
Jia, Shuncheng
Li, Zhong-Zhi
Zhao, Xuanle
Bai, Ye
Chen, Feilong
Tian, Qi
Zhang, Tielin
author_facet Zhang, Duzhen
Ren, Yong
Cong, Wei
Zheng, Junhao
Su, Qiaoyi
Jia, Shuncheng
Li, Zhong-Zhi
Zhao, Xuanle
Bai, Ye
Chen, Feilong
Tian, Qi
Zhang, Tielin
contents Continual Semantic Segmentation (CSS) seeks to incrementally learn to segment novel classes while preserving knowledge of previously encountered ones. Recent advancements in CSS have been largely driven by the adoption of Pre-trained Vision Models (PVMs) as backbones. Among existing strategies, Direct Fine-Tuning (DFT), which sequentially fine-tunes the model across classes, remains the most straightforward approach. Prior work often regards DFT as a performance lower bound due to its presumed vulnerability to severe catastrophic forgetting, leading to the development of numerous complex mitigation techniques. However, we contend that this prevailing assumption is flawed. In this paper, we systematically revisit forgetting in DFT across two standard benchmarks, Pascal VOC 2012 and ADE20K, under eight CSS settings using two representative PVM backbones: ResNet101 and Swin-B. Through a detailed probing analysis, our findings reveal that existing methods significantly underestimate the inherent anti-forgetting capabilities of PVMs. Even under DFT, PVMs retain previously learned knowledge with minimal forgetting. Further investigation of the feature space indicates that the observed forgetting primarily arises from the classifier's drift away from the PVM, rather than from degradation of the backbone representations. Based on this insight, we propose DFT*, a simple yet effective enhancement to DFT that incorporates strategies such as freezing the PVM backbone and previously learned classifiers, as well as pre-allocating future classifiers. Extensive experiments show that DFT* consistently achieves competitive or superior performance compared to sixteen state-of-the-art CSS methods, while requiring substantially fewer trainable parameters and less training time.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
Zhang, Duzhen
Ren, Yong
Cong, Wei
Zheng, Junhao
Su, Qiaoyi
Jia, Shuncheng
Li, Zhong-Zhi
Zhao, Xuanle
Bai, Ye
Chen, Feilong
Tian, Qi
Zhang, Tielin
Computer Vision and Pattern Recognition
Continual Semantic Segmentation (CSS) seeks to incrementally learn to segment novel classes while preserving knowledge of previously encountered ones. Recent advancements in CSS have been largely driven by the adoption of Pre-trained Vision Models (PVMs) as backbones. Among existing strategies, Direct Fine-Tuning (DFT), which sequentially fine-tunes the model across classes, remains the most straightforward approach. Prior work often regards DFT as a performance lower bound due to its presumed vulnerability to severe catastrophic forgetting, leading to the development of numerous complex mitigation techniques. However, we contend that this prevailing assumption is flawed. In this paper, we systematically revisit forgetting in DFT across two standard benchmarks, Pascal VOC 2012 and ADE20K, under eight CSS settings using two representative PVM backbones: ResNet101 and Swin-B. Through a detailed probing analysis, our findings reveal that existing methods significantly underestimate the inherent anti-forgetting capabilities of PVMs. Even under DFT, PVMs retain previously learned knowledge with minimal forgetting. Further investigation of the feature space indicates that the observed forgetting primarily arises from the classifier's drift away from the PVM, rather than from degradation of the backbone representations. Based on this insight, we propose DFT*, a simple yet effective enhancement to DFT that incorporates strategies such as freezing the PVM backbone and previously learned classifiers, as well as pre-allocating future classifiers. Extensive experiments show that DFT* consistently achieves competitive or superior performance compared to sixteen state-of-the-art CSS methods, while requiring substantially fewer trainable parameters and less training time.
title Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.04267