Saved in:
Bibliographic Details
Main Authors: Cai, Chaoxiang, Yang, Longrong, Weng, Minghe, Li, Xuewei, Qin, Zequn, Li, Xi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.01351
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914438277234688
author Cai, Chaoxiang
Yang, Longrong
Weng, Minghe
Li, Xuewei
Qin, Zequn
Li, Xi
author_facet Cai, Chaoxiang
Yang, Longrong
Weng, Minghe
Li, Xuewei
Qin, Zequn
Li, Xi
contents The mixture-of-experts (MoE) architecture, which replaces dense networks with sparse ones, has attracted significant attention in large vision-language models (LVLMs) for achieving comparable performance while activating far fewer parameters. Existing MoE architectures for LVLMs primarily focus on token-to-expert routing (TER), encouraging different experts to specialize in processing specific tokens. However, these methods typically rely on the load balancing mechanism, neglecting the inherent distributional differences between vision and language modalities. To address this limitation, we propose the Long-Tailed Distribution-aware Router (LTDR) for vision-language TER, which tackles two key challenges: (1) Modality-specific distribution-aware routing. We observe that language TER generally follows a relatively uniform distribution, whereas vision TER exhibits a long-tailed distribution. This modality discrepancy motivates the design of specialized routing strategies for each modality. (2) Vision-specific dynamic expert activation. Recognizing the importance of high-information vision tail tokens, we introduce a data-augmentation-inspired strategy that increases the number of activated experts, ensuring sufficient learning for these rare but informative tokens. On vision-language and vision benchmarks, our approach achieves consistent improvements, boosting performance by 1.2% / 2.1% on vision-language and 1.6% on vision benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
Cai, Chaoxiang
Yang, Longrong
Weng, Minghe
Li, Xuewei
Qin, Zequn
Li, Xi
Computer Vision and Pattern Recognition
The mixture-of-experts (MoE) architecture, which replaces dense networks with sparse ones, has attracted significant attention in large vision-language models (LVLMs) for achieving comparable performance while activating far fewer parameters. Existing MoE architectures for LVLMs primarily focus on token-to-expert routing (TER), encouraging different experts to specialize in processing specific tokens. However, these methods typically rely on the load balancing mechanism, neglecting the inherent distributional differences between vision and language modalities. To address this limitation, we propose the Long-Tailed Distribution-aware Router (LTDR) for vision-language TER, which tackles two key challenges: (1) Modality-specific distribution-aware routing. We observe that language TER generally follows a relatively uniform distribution, whereas vision TER exhibits a long-tailed distribution. This modality discrepancy motivates the design of specialized routing strategies for each modality. (2) Vision-specific dynamic expert activation. Recognizing the importance of high-information vision tail tokens, we introduce a data-augmentation-inspired strategy that increases the number of activated experts, ensuring sufficient learning for these rare but informative tokens. On vision-language and vision benchmarks, our approach achieves consistent improvements, boosting performance by 1.2% / 2.1% on vision-language and 1.6% on vision benchmarks.
title Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01351