Modality-Inconsistent Continual Learning of Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pian, Weiguo, Deng, Shijian, Mo, Shentong, Liu, Mingrui, Guo, Yunhui, Tian, Yapeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910209973157888
author Pian, Weiguo
Deng, Shijian
Mo, Shentong
Liu, Mingrui
Guo, Yunhui
Tian, Yapeng
author_facet Pian, Weiguo
Deng, Shijian
Mo, Shentong
Liu, Mingrui
Guo, Yunhui
Tian, Yapeng
contents In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with inconsistent modalities (image, audio, or video) and varying task types (captioning or question-answering). Unlike existing vision-only or modality-incremental settings, MICL combines modality and task type shifts, both of which drive catastrophic forgetting. To address these challenges, we propose MoInCL, which employs a Pseudo Targets Generation Module to mitigate forgetting caused by task type shifts in previously seen modalities. It also incorporates Instruction-based Knowledge Distillation to preserve the model's ability to handle previously learned modalities when new ones are introduced. We benchmark MICL using a total of six tasks and conduct experiments to validate the effectiveness of our MoInCL. The experimental results highlight the superiority of MoInCL, showing significant improvements over representative and state-of-the-art continual learning baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13050
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Modality-Inconsistent Continual Learning of Multimodal Large Language Models
Pian, Weiguo
Deng, Shijian
Mo, Shentong
Liu, Mingrui
Guo, Yunhui
Tian, Yapeng
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with inconsistent modalities (image, audio, or video) and varying task types (captioning or question-answering). Unlike existing vision-only or modality-incremental settings, MICL combines modality and task type shifts, both of which drive catastrophic forgetting. To address these challenges, we propose MoInCL, which employs a Pseudo Targets Generation Module to mitigate forgetting caused by task type shifts in previously seen modalities. It also incorporates Instruction-based Knowledge Distillation to preserve the model's ability to handle previously learned modalities when new ones are introduced. We benchmark MICL using a total of six tasks and conduct experiments to validate the effectiveness of our MoInCL. The experimental results highlight the superiority of MoInCL, showing significant improvements over representative and state-of-the-art continual learning baselines.
title Modality-Inconsistent Continual Learning of Multimodal Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.13050