Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Longrong, Shen, Dong, Cai, Chaoxiang, Yang, Fan, Gao, Tingting, Zhang, Di, Li, Xi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913736827076608
author Yang, Longrong
Shen, Dong
Cai, Chaoxiang
Yang, Fan
Gao, Tingting
Zhang, Di
Li, Xi
author_facet Yang, Longrong
Shen, Dong
Cai, Chaoxiang
Yang, Fan
Gao, Tingting
Zhang, Di
Li, Xi
contents The Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-Language Models (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Existing MoE methods in LVLM encourage different experts to specialize in different tokens, and they usually employ a router to predict the routing of each token. However, the router is not optimized concerning distinct parameter optimization directions generated from tokens within an expert. This may lead to severe interference between tokens within an expert. To address this problem, we propose to use the token-level gradient analysis to Solving Token Gradient Conflict (STGC) in this paper. Specifically, we first use token-level gradients to identify conflicting tokens in experts. After that, we add a regularization loss tailored to encourage conflicting tokens routing from their current experts to other experts, for reducing interference between tokens within an expert. Our method can serve as a plug-in for diverse LVLM methods, and extensive experimental results demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/STGC.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19905
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
Yang, Longrong
Shen, Dong
Cai, Chaoxiang
Yang, Fan
Gao, Tingting
Zhang, Di
Li, Xi
Computer Vision and Pattern Recognition
The Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-Language Models (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Existing MoE methods in LVLM encourage different experts to specialize in different tokens, and they usually employ a router to predict the routing of each token. However, the router is not optimized concerning distinct parameter optimization directions generated from tokens within an expert. This may lead to severe interference between tokens within an expert. To address this problem, we propose to use the token-level gradient analysis to Solving Token Gradient Conflict (STGC) in this paper. Specifically, we first use token-level gradients to identify conflicting tokens in experts. After that, we add a regularization loss tailored to encourage conflicting tokens routing from their current experts to other experts, for reducing interference between tokens within an expert. Our method can serve as a plug-in for diverse LVLM methods, and extensive experimental results demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/STGC.
title Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.19905