Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Haiyang, Zeng, Fanhu, Zhu, Fei, Wang, Jiayi, Wang, Xukai, Zhou, Jingang, Zhao, Hongbo, Liu, Wenzhuo, Ma, Shijie, Wang, Da-Han, Zhang, Xu-Yao, Liu, Cheng-Lin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912551021838336
author Guo, Haiyang
Zeng, Fanhu
Zhu, Fei
Wang, Jiayi
Wang, Xukai
Zhou, Jingang
Zhao, Hongbo
Liu, Wenzhuo
Ma, Shijie
Wang, Da-Han
Zhang, Xu-Yao
Liu, Cheng-Lin
author_facet Guo, Haiyang
Zeng, Fanhu
Zhu, Fei
Wang, Jiayi
Wang, Xukai
Zhou, Jingang
Zhao, Hongbo
Liu, Wenzhuo
Ma, Shijie
Wang, Da-Han
Zhang, Xu-Yao
Liu, Cheng-Lin
contents The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specific domains. However, these models are fundamentally constrained by \emph{catastrophic forgetting}, \ie~a persistent challenge where models experience performance degradation on previously learned tasks when adapting to new tasks. To address this practical limitation, numerous approaches have been proposed to enhance the adaptability and scalability of generative AI in real-world applications. In this work, we present a comprehensive survey of continual learning methods for mainstream generative AI models, encompassing large language models, multimodal large language models, vision-language-action models, and diffusion models. Drawing inspiration from the memory mechanisms of the human brain, we systematically categorize these approaches into three paradigms: architecture-based, regularization-based, and replay-based methods, while elucidating their underlying methodologies and motivations. We further analyze continual learning setups for different generative models, including training objectives, benchmarks, and core backbones, thereby providing deeper insights into the field. The project page of this paper is available at https://github.com/Ghy0501/Awesome-Continual-Learning-in-Generative-Models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
Guo, Haiyang
Zeng, Fanhu
Zhu, Fei
Wang, Jiayi
Wang, Xukai
Zhou, Jingang
Zhao, Hongbo
Liu, Wenzhuo
Ma, Shijie
Wang, Da-Han
Zhang, Xu-Yao
Liu, Cheng-Lin
Machine Learning
Computer Vision and Pattern Recognition
The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specific domains. However, these models are fundamentally constrained by \emph{catastrophic forgetting}, \ie~a persistent challenge where models experience performance degradation on previously learned tasks when adapting to new tasks. To address this practical limitation, numerous approaches have been proposed to enhance the adaptability and scalability of generative AI in real-world applications. In this work, we present a comprehensive survey of continual learning methods for mainstream generative AI models, encompassing large language models, multimodal large language models, vision-language-action models, and diffusion models. Drawing inspiration from the memory mechanisms of the human brain, we systematically categorize these approaches into three paradigms: architecture-based, regularization-based, and replay-based methods, while elucidating their underlying methodologies and motivations. We further analyze continual learning setups for different generative models, including training objectives, benchmarks, and core backbones, thereby providing deeper insights into the field. The project page of this paper is available at https://github.com/Ghy0501/Awesome-Continual-Learning-in-Generative-Models.
title Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13045