GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Tiantian, Liu, Yuyang, Yang, Shuo, Hong, Qiuhe, Tian, YongHong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915411018121216
author Peng, Tiantian
Liu, Yuyang
Yang, Shuo
Hong, Qiuhe
Tian, YongHong
author_facet Peng, Tiantian
Liu, Yuyang
Yang, Shuo
Hong, Qiuhe
Tian, YongHong
contents Contrastive Language-Image Pretraining has demonstrated remarkable zero-shot generalization by aligning visual and textual modalities in a shared embedding space. However, when continuously fine-tuned on diverse tasks, CLIP suffers from catastrophic forgetting and degradation of its embedding alignment, undermining its zero-shot capabilities. In this work, we propose Gradient Null Space Projection (GNSP), an efficient continual learning method that projects task-specific gradients onto the null space of previously learned knowledge. This orthogonal projection mathematically prevents interference with previous tasks without relying on rehearsal or architectural modification. Furthermore, to preserve the inherent generalization property of CLIP, we introduce knowledge distillation and combine it with a modality alignment preservation loss inspired by CLIP pre-training to stabilize the structure of the multimodal embedding space during fine-tuning. On the MTIL benchmark consisting of 11 tasks, our method achieved SOTA performance on both the Average and Last key metrics. More importantly, experiments show that our method successfully maintains the original modality gap and cross-modal retrieval performance of CLIP, confirming its effectiveness in maintaining a robust visual-language space throughout the continual learning process.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning
Peng, Tiantian
Liu, Yuyang
Yang, Shuo
Hong, Qiuhe
Tian, YongHong
Machine Learning
Computer Vision and Pattern Recognition
Contrastive Language-Image Pretraining has demonstrated remarkable zero-shot generalization by aligning visual and textual modalities in a shared embedding space. However, when continuously fine-tuned on diverse tasks, CLIP suffers from catastrophic forgetting and degradation of its embedding alignment, undermining its zero-shot capabilities. In this work, we propose Gradient Null Space Projection (GNSP), an efficient continual learning method that projects task-specific gradients onto the null space of previously learned knowledge. This orthogonal projection mathematically prevents interference with previous tasks without relying on rehearsal or architectural modification. Furthermore, to preserve the inherent generalization property of CLIP, we introduce knowledge distillation and combine it with a modality alignment preservation loss inspired by CLIP pre-training to stabilize the structure of the multimodal embedding space during fine-tuning. On the MTIL benchmark consisting of 11 tasks, our method achieved SOTA performance on both the Average and Last key metrics. More importantly, experiments show that our method successfully maintains the original modality gap and cross-modal retrieval performance of CLIP, confirming its effectiveness in maintaining a robust visual-language space throughout the continual learning process.
title GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19839