MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908739492118528 |
|---|---|
| author | Guo, Haiyang Zhu, Fei Zhao, Hongbo Zeng, Fanhu Liu, Wenzhuo Ma, Shijie Wang, Da-Han Zhang, Xu-Yao |
| author_facet | Guo, Haiyang Zhu, Fei Zhao, Hongbo Zeng, Fanhu Liu, Wenzhuo Ma, Shijie Wang, Da-Han Zhang, Xu-Yao |
| contents | Continual learning enables AI systems to acquire new knowledge while retaining previously learned information. While traditional unimodal methods have made progress, the rise of Multimodal Large Language Models (MLLMs) brings new challenges in Multimodal Continual Learning (MCL), where models are expected to address both catastrophic forgetting and cross-modal coordination. To advance research in this area, we present MCITlib, a comprehensive library for Multimodal Continual Instruction Tuning. MCITlib currently implements 8 representative algorithms and conducts evaluations on 3 benchmarks under 2 backbone models. The library will be continuously updated to support future developments in MCL. The codebase is released at https://github.com/Ghy0501/MCITlib. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_07307 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark Guo, Haiyang Zhu, Fei Zhao, Hongbo Zeng, Fanhu Liu, Wenzhuo Ma, Shijie Wang, Da-Han Zhang, Xu-Yao Computer Vision and Pattern Recognition Artificial Intelligence Continual learning enables AI systems to acquire new knowledge while retaining previously learned information. While traditional unimodal methods have made progress, the rise of Multimodal Large Language Models (MLLMs) brings new challenges in Multimodal Continual Learning (MCL), where models are expected to address both catastrophic forgetting and cross-modal coordination. To advance research in this area, we present MCITlib, a comprehensive library for Multimodal Continual Instruction Tuning. MCITlib currently implements 8 representative algorithms and conducts evaluations on 3 benchmarks under 2 backbone models. The library will be continuously updated to support future developments in MCL. The codebase is released at https://github.com/Ghy0501/MCITlib. |
| title | MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2508.07307 |