Continual Learning in Vision-Language Models via Aligned Model Merging

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sokar, Ghada, Dziugaite, Gintare Karolina, Arnab, Anurag, Iscen, Ahmet, Castro, Pablo Samuel, Schmid, Cordelia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910985821880320
author Sokar, Ghada
Dziugaite, Gintare Karolina
Arnab, Anurag
Iscen, Ahmet
Castro, Pablo Samuel
Schmid, Cordelia
author_facet Sokar, Ghada
Dziugaite, Gintare Karolina
Arnab, Anurag
Iscen, Ahmet
Castro, Pablo Samuel
Schmid, Cordelia
contents Continual learning is conventionally tackled through sequential fine-tuning, a process that, while enabling adaptation, inherently favors plasticity over the stability needed to retain prior knowledge. While existing approaches attempt to mitigate catastrophic forgetting, a bias towards recent tasks persists as they build upon this sequential nature. In this work we present a new perspective based on model merging to maintain stability while still retaining plasticity. Rather than just sequentially updating the model weights, we propose merging newly trained task parameters with previously learned ones, promoting a better balance. To maximize the effectiveness of the merging process, we propose a simple mechanism that promotes learning aligned weights with previous ones, thereby avoiding interference when merging. We evaluate this approach on large Vision-Language Models (VLMs), and demonstrate its effectiveness in reducing forgetting, increasing robustness to various task orders and similarities, and improving generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continual Learning in Vision-Language Models via Aligned Model Merging
Sokar, Ghada
Dziugaite, Gintare Karolina
Arnab, Anurag
Iscen, Ahmet
Castro, Pablo Samuel
Schmid, Cordelia
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Continual learning is conventionally tackled through sequential fine-tuning, a process that, while enabling adaptation, inherently favors plasticity over the stability needed to retain prior knowledge. While existing approaches attempt to mitigate catastrophic forgetting, a bias towards recent tasks persists as they build upon this sequential nature. In this work we present a new perspective based on model merging to maintain stability while still retaining plasticity. Rather than just sequentially updating the model weights, we propose merging newly trained task parameters with previously learned ones, promoting a better balance. To maximize the effectiveness of the merging process, we propose a simple mechanism that promotes learning aligned weights with previous ones, thereby avoiding interference when merging. We evaluate this approach on large Vision-Language Models (VLMs), and demonstrate its effectiveness in reducing forgetting, increasing robustness to various task orders and similarities, and improving generalization.
title Continual Learning in Vision-Language Models via Aligned Model Merging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.03189