Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gou, Yunhao, Liu, Zhili, Chen, Kai, Hong, Lanqing, Xu, Hang, Li, Aoxue, Yeung, Dit-Yan, Kwok, James T., Zhang, Yu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916311171334144
author Gou, Yunhao
Liu, Zhili
Chen, Kai
Hong, Lanqing
Xu, Hang
Li, Aoxue
Yeung, Dit-Yan
Kwok, James T.
Zhang, Yu
author_facet Gou, Yunhao
Liu, Zhili
Chen, Kai
Hong, Lanqing
Xu, Hang
Li, Aoxue
Yeung, Dit-Yan
Kwok, James T.
Zhang, Yu
contents Instruction tuning of Large Vision-language Models (LVLMs) has revolutionized the development of versatile models with zero-shot generalization across a wide range of downstream vision-language tasks. However, the diversity of training tasks of different sources and formats would lead to inevitable task conflicts, where different tasks conflict for the same set of model parameters, resulting in sub-optimal instruction-following abilities. To address that, we propose the Mixture of Cluster-conditional LoRA Experts (MoCLE), a novel Mixture of Experts (MoE) architecture designed to activate the task-customized model parameters based on the instruction clusters. A separate universal expert is further incorporated to improve generalization capabilities of MoCLE for novel instructions. Extensive experiments on InstructBLIP and LLaVA demonstrate the effectiveness of MoCLE.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12379
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
Gou, Yunhao
Liu, Zhili
Chen, Kai
Hong, Lanqing
Xu, Hang
Li, Aoxue
Yeung, Dit-Yan
Kwok, James T.
Zhang, Yu
Computer Vision and Pattern Recognition
Instruction tuning of Large Vision-language Models (LVLMs) has revolutionized the development of versatile models with zero-shot generalization across a wide range of downstream vision-language tasks. However, the diversity of training tasks of different sources and formats would lead to inevitable task conflicts, where different tasks conflict for the same set of model parameters, resulting in sub-optimal instruction-following abilities. To address that, we propose the Mixture of Cluster-conditional LoRA Experts (MoCLE), a novel Mixture of Experts (MoE) architecture designed to activate the task-customized model parameters based on the instruction clusters. A separate universal expert is further incorporated to improve generalization capabilities of MoCLE for novel instructions. Extensive experiments on InstructBLIP and LLaVA demonstrate the effectiveness of MoCLE.
title Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.12379