CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiachen, Wang, Xinyao, Zhu, Sijie, Kuo, Chia-Wen, Xu, Lu, Chen, Fan, Jain, Jitesh, Shi, Humphrey, Wen, Longyin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909196899844096
author Li, Jiachen
Wang, Xinyao
Zhu, Sijie
Kuo, Chia-Wen
Xu, Lu
Chen, Fan
Jain, Jitesh
Shi, Humphrey
Wen, Longyin
author_facet Li, Jiachen
Wang, Xinyao
Zhu, Sijie
Kuo, Chia-Wen
Xu, Lu
Chen, Fan
Jain, Jitesh
Shi, Humphrey
Wen, Longyin
contents Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of improving model capabilities from the vision side. Inspired by the successful applications of Mixture-of-Experts (MoE) in LLMs, which improves model scalability during training while keeping inference costs similar to those of smaller models, we propose CuMo. CuMo incorporates Co-upcycled Top-K sparsely-gated Mixture-of-experts blocks into both the vision encoder and the MLP connector, thereby enhancing the multimodal LLMs with minimal additional activated parameters during inference. CuMo first pre-trains the MLP blocks and then initializes each expert in the MoE block from the pre-trained MLP block during the visual instruction tuning stage. Auxiliary losses are used to ensure a balanced loading of experts. CuMo outperforms state-of-the-art multimodal LLMs across various VQA and visual-instruction-following benchmarks using models within each model size group, all while training exclusively on open-sourced datasets. The code and model weights for CuMo are open-sourced at https://github.com/SHI-Labs/CuMo.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05949
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
Li, Jiachen
Wang, Xinyao
Zhu, Sijie
Kuo, Chia-Wen
Xu, Lu
Chen, Fan
Jain, Jitesh
Shi, Humphrey
Wen, Longyin
Computer Vision and Pattern Recognition
Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of improving model capabilities from the vision side. Inspired by the successful applications of Mixture-of-Experts (MoE) in LLMs, which improves model scalability during training while keeping inference costs similar to those of smaller models, we propose CuMo. CuMo incorporates Co-upcycled Top-K sparsely-gated Mixture-of-experts blocks into both the vision encoder and the MLP connector, thereby enhancing the multimodal LLMs with minimal additional activated parameters during inference. CuMo first pre-trains the MLP blocks and then initializes each expert in the MoE block from the pre-trained MLP block during the visual instruction tuning stage. Auxiliary losses are used to ensure a balanced loading of experts. CuMo outperforms state-of-the-art multimodal LLMs across various VQA and visual-instruction-following benchmarks using models within each model size group, all while training exclusively on open-sourced datasets. The code and model weights for CuMo are open-sourced at https://github.com/SHI-Labs/CuMo.
title CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.05949