iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Xuezhi, Zhou, Dongbo, Guo, Wang, Wang, Zeyuan, Li, Ziyu, Zhou, Gaozhi, Li, Xian, Zhao, Ling, Yang, Wentao, Tao, Chao, Li, Haifeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913144437211136
author Cui, Xuezhi
Zhou, Dongbo
Guo, Wang
Wang, Zeyuan
Li, Ziyu
Zhou, Gaozhi
Li, Xian
Zhao, Ling
Yang, Wentao
Tao, Chao
Li, Haifeng
author_facet Cui, Xuezhi
Zhou, Dongbo
Guo, Wang
Wang, Zeyuan
Li, Ziyu
Zhou, Gaozhi
Li, Xian
Zhao, Ling
Yang, Wentao
Tao, Chao
Li, Haifeng
contents Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning isolated modules per task leads to parameter explosion. Conversely, recent similarity-driven sharing mechanisms falsely equate superficial visual similarity with underlying alignment consistency. This fundamental mismatch triggers severe negative transfer between visually similar but logically distinct tasks and fails to exploit alignment reuse across visually diverse ones. We argue thatalignment sharing is fundamentally a geometric problem of overlapping optimization trajectories within shared low-rank subspaces. Grounded in this insight, we propose iGSP, a novel framework that achieves efficient adaptation via implicit gradient subspace projection. Leveraging the early convergence of MoE routers to establish the subspace basis, iGSP bifurcates the adaptation process into two phases. First, the Subspace Identification phase introduces candidate experts via basis pre-expansion, applies a novel subspace-constrained regularization to implicitly project new task gradients onto the historical subspace, and precisely prunes redundant dimensions by treating routing probabilities as gradient flow indicators, ultimately to maximize knowledge reuse. Second, the Orthogonal Subspace Fine-Tuning phase fixes this structural basis and removes the regularization to rapidly fit the task-specific residual loss. Extensive experiments on the MTIL benchmark demonstrate that iGSP achieves state-of-the-art accuracy while significantly improving training efficiency, reducing the average trainable parameters by 42.7\% compared to current SOTA methods, and decreasing the final total parameters by 86.9\% relative to counterparts. The source code is available at https://github.com/GeoX-Lab/iGSP.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19301
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models
Cui, Xuezhi
Zhou, Dongbo
Guo, Wang
Wang, Zeyuan
Li, Ziyu
Zhou, Gaozhi
Li, Xian
Zhao, Ling
Yang, Wentao
Tao, Chao
Li, Haifeng
Computer Vision and Pattern Recognition
Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning isolated modules per task leads to parameter explosion. Conversely, recent similarity-driven sharing mechanisms falsely equate superficial visual similarity with underlying alignment consistency. This fundamental mismatch triggers severe negative transfer between visually similar but logically distinct tasks and fails to exploit alignment reuse across visually diverse ones. We argue thatalignment sharing is fundamentally a geometric problem of overlapping optimization trajectories within shared low-rank subspaces. Grounded in this insight, we propose iGSP, a novel framework that achieves efficient adaptation via implicit gradient subspace projection. Leveraging the early convergence of MoE routers to establish the subspace basis, iGSP bifurcates the adaptation process into two phases. First, the Subspace Identification phase introduces candidate experts via basis pre-expansion, applies a novel subspace-constrained regularization to implicitly project new task gradients onto the historical subspace, and precisely prunes redundant dimensions by treating routing probabilities as gradient flow indicators, ultimately to maximize knowledge reuse. Second, the Orthogonal Subspace Fine-Tuning phase fixes this structural basis and removes the regularization to rapidly fit the task-specific residual loss. Extensive experiments on the MTIL benchmark demonstrate that iGSP achieves state-of-the-art accuracy while significantly improving training efficiency, reducing the average trainable parameters by 42.7\% compared to current SOTA methods, and decreasing the final total parameters by 86.9\% relative to counterparts. The source code is available at https://github.com/GeoX-Lab/iGSP.
title iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.19301