p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jun, Meng, Desen, Zhang, Zhengming, Huang, Zhenpeng, Wu, Tao, Wang, Limin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912522414587904
author Zhang, Jun
Meng, Desen
Zhang, Zhengming
Huang, Zhenpeng
Wu, Tao
Wang, Limin
author_facet Zhang, Jun
Meng, Desen
Zhang, Zhengming
Huang, Zhenpeng
Wu, Tao
Wang, Limin
contents Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while maintaining model performance. The majority of computation in MLLMs stems from the overwhelming volume of vision tokens processed by the transformer-based LLM. Accordingly, we leverage the Mixture-of-Depths (MoD) mechanism, where each LLM layer selects essential vision tokens to process while skipping redundant ones. However, integrating MoD into MLLMs is non-trivial. To address the challenges of training and inference stability as well as limited training data, we adapt the MoD module with two novel designs: tanh-gated weight normalization (TanhNorm) and symmetric token reweighting (STRing). Moreover, we observe that vision tokens exhibit higher redundancy in deeper layers and thus design a progressive ratio decay (PRD) strategy, which gradually reduces the token retention ratio layer by layer, employing a shifted cosine schedule. This crucial design fully unleashes the potential of MoD, significantly boosting the efficiency and performance of our models. Extensive experiments on two baseline models across 15 benchmarks show that our model matches or even surpasses the performance of corresponding baselines, while requiring only 55.6% TFLOPs and 53.7% KV cache storage during inference, and 77.7% GPU hours during training.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04449
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
Zhang, Jun
Meng, Desen
Zhang, Zhengming
Huang, Zhenpeng
Wu, Tao
Wang, Limin
Computer Vision and Pattern Recognition
Computation and Language
Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while maintaining model performance. The majority of computation in MLLMs stems from the overwhelming volume of vision tokens processed by the transformer-based LLM. Accordingly, we leverage the Mixture-of-Depths (MoD) mechanism, where each LLM layer selects essential vision tokens to process while skipping redundant ones. However, integrating MoD into MLLMs is non-trivial. To address the challenges of training and inference stability as well as limited training data, we adapt the MoD module with two novel designs: tanh-gated weight normalization (TanhNorm) and symmetric token reweighting (STRing). Moreover, we observe that vision tokens exhibit higher redundancy in deeper layers and thus design a progressive ratio decay (PRD) strategy, which gradually reduces the token retention ratio layer by layer, employing a shifted cosine schedule. This crucial design fully unleashes the potential of MoD, significantly boosting the efficiency and performance of our models. Extensive experiments on two baseline models across 15 benchmarks show that our model matches or even surpasses the performance of corresponding baselines, while requiring only 55.6% TFLOPs and 53.7% KV cache storage during inference, and 77.7% GPU hours during training.
title p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.04449