Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Durrani, Hamza Ahmed, Durrani, Rafay Suleman
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910214161170432
author Durrani, Hamza Ahmed
Durrani, Rafay Suleman
author_facet Durrani, Hamza Ahmed
Durrani, Rafay Suleman
contents Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalities. These models excel at tasks like image captioning, visual question answering, and cross-modal retrieval. However, they face catastrophic forgetting when learning new tasks sequentially, particularly challenging in multi-modal settings where preserving cross-modal alignments adds complexity to the learning process. This paper presents a comprehensive continual learning framework for LVLMs that combines enhanced Elastic Weight Consolidation (EWC) with parameter-efficient fine-tuning techniques. We integrate multi-modal Fisher Information Matrix calculation, consistency preservation across modalities, and adaptive regularization that considers dependencies across visual and textual encoders. The framework achieves a 78% reduction in forgetting rates relative to naive sequential training approaches through extensive evaluation testing. The framework also preserves alignment between modalities during sequential learning with only 15% additional computational cost. This work advances the state of the art in lifelong learning for multi-modal AI systems, with direct applications to autonomous driving, intelligent robotic assistants, and adaptive robotic systems that must continuously learn in dynamic real-world environments.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
Durrani, Hamza Ahmed
Durrani, Rafay Suleman
Robotics
I.2.10; I.4.8; I.2.6
Large language-vision models (LVLMs) such as CLIP, Flamingo, and BLIP have revolutionized AI by enabling understanding across textual and visual modalities. These models excel at tasks like image captioning, visual question answering, and cross-modal retrieval. However, they face catastrophic forgetting when learning new tasks sequentially, particularly challenging in multi-modal settings where preserving cross-modal alignments adds complexity to the learning process. This paper presents a comprehensive continual learning framework for LVLMs that combines enhanced Elastic Weight Consolidation (EWC) with parameter-efficient fine-tuning techniques. We integrate multi-modal Fisher Information Matrix calculation, consistency preservation across modalities, and adaptive regularization that considers dependencies across visual and textual encoders. The framework achieves a 78% reduction in forgetting rates relative to naive sequential training approaches through extensive evaluation testing. The framework also preserves alignment between modalities during sequential learning with only 15% additional computational cost. This work advances the state of the art in lifelong learning for multi-modal AI systems, with direct applications to autonomous driving, intelligent robotic assistants, and adaptive robotic systems that must continuously learn in dynamic real-world environments.
title Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
topic Robotics
I.2.10; I.4.8; I.2.6
url https://arxiv.org/abs/2605.12789