Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Wenke, Liang, Jian, Guo, Xianda, Fang, Yiyang, Wan, Guancheng, Rong, Xuankun, Wen, Chi, Shi, Zekun, Li, Qingyun, Zhu, Didi, Ma, Yanbiao, Liang, Ke, Yang, Bin, Li, He, Shao, Jiawei, Ye, Mang, Du, Bo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916645265473536
author Huang, Wenke
Liang, Jian
Guo, Xianda
Fang, Yiyang
Wan, Guancheng
Rong, Xuankun
Wen, Chi
Shi, Zekun
Li, Qingyun
Zhu, Didi
Ma, Yanbiao
Liang, Ke
Yang, Bin
Li, He
Shao, Jiawei
Ye, Mang
Du, Bo
author_facet Huang, Wenke
Liang, Jian
Guo, Xianda
Fang, Yiyang
Wan, Guancheng
Rong, Xuankun
Wen, Chi
Shi, Zekun
Li, Qingyun
Zhu, Didi
Ma, Yanbiao
Liang, Ke
Yang, Bin
Li, He
Shao, Jiawei
Ye, Mang
Du, Bo
contents Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs appears limited performance on special applications. But tuning MLLMs for downstream tasks encounters two key challenges: Task-Expert Specialization, where distribution shifts between pre-training and target datasets constrain target performance, and Open-World Stabilization, where catastrophic forgetting erases the model general knowledge. In this work, we systematically review recent advancements in MLLM tuning methodologies, classifying them into three paradigms: (I) Selective Tuning, (II) Additive Tuning, and (III) Reparameterization Tuning. Furthermore, we benchmark these tuning strategies across popular MLLM architectures and diverse downstream tasks to establish standardized evaluation analysis and systematic tuning principles. Finally, we highlight several open challenges in this domain and propose future research directions. To facilitate ongoing progress in this rapidly evolving field, we provide a public repository that continuously tracks developments: https://github.com/WenkeHuang/Awesome-MLLM-Tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
Huang, Wenke
Liang, Jian
Guo, Xianda
Fang, Yiyang
Wan, Guancheng
Rong, Xuankun
Wen, Chi
Shi, Zekun
Li, Qingyun
Zhu, Didi
Ma, Yanbiao
Liang, Ke
Yang, Bin
Li, He
Shao, Jiawei
Ye, Mang
Du, Bo
Computation and Language
Artificial Intelligence
Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs appears limited performance on special applications. But tuning MLLMs for downstream tasks encounters two key challenges: Task-Expert Specialization, where distribution shifts between pre-training and target datasets constrain target performance, and Open-World Stabilization, where catastrophic forgetting erases the model general knowledge. In this work, we systematically review recent advancements in MLLM tuning methodologies, classifying them into three paradigms: (I) Selective Tuning, (II) Additive Tuning, and (III) Reparameterization Tuning. Furthermore, we benchmark these tuning strategies across popular MLLM architectures and diverse downstream tasks to establish standardized evaluation analysis and systematic tuning principles. Finally, we highlight several open challenges in this domain and propose future research directions. To facilitate ongoing progress in this rapidly evolving field, we provide a public repository that continuously tracks developments: https://github.com/WenkeHuang/Awesome-MLLM-Tuning.
title Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.04543