Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Longrong, Shen, Dong, Cai, Chaoxiang, Yang, Fan, Gao, Tingting, Zhang, Di, Li, Xi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
di: Cai, Chaoxiang, et al.
Pubblicazione: (2025)
di: Cai, Chaoxiang, et al.
Pubblicazione: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
di: Jia, Chenwei, et al.
Pubblicazione: (2026)
di: Jia, Chenwei, et al.
Pubblicazione: (2026)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
di: Cao, Sihan, et al.
Pubblicazione: (2026)
di: Cao, Sihan, et al.
Pubblicazione: (2026)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
Generalizable Multispectral Land Cover Classification via Frequency-Aware Mixture of Low-Rank Token Experts
di: Chen, Xi, et al.
Pubblicazione: (2025)
di: Chen, Xi, et al.
Pubblicazione: (2025)
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
di: Liu, Jiaqi, et al.
Pubblicazione: (2025)
di: Liu, Jiaqi, et al.
Pubblicazione: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
di: Shen, Leyang, et al.
Pubblicazione: (2024)
di: Shen, Leyang, et al.
Pubblicazione: (2024)
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
di: He, Xin, et al.
Pubblicazione: (2025)
di: He, Xin, et al.
Pubblicazione: (2025)
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
di: Dong, Shaoqi, et al.
Pubblicazione: (2025)
di: Dong, Shaoqi, et al.
Pubblicazione: (2025)
Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
di: Zhang, Jiaxing, et al.
Pubblicazione: (2025)
di: Zhang, Jiaxing, et al.
Pubblicazione: (2025)
DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models
di: Qin, Mengxin, et al.
Pubblicazione: (2026)
di: Qin, Mengxin, et al.
Pubblicazione: (2026)
EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
di: Feng, Ze, et al.
Pubblicazione: (2025)
di: Feng, Ze, et al.
Pubblicazione: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
di: Wen, Zichen, et al.
Pubblicazione: (2025)
di: Wen, Zichen, et al.
Pubblicazione: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts
di: Chen, Xi, et al.
Pubblicazione: (2026)
di: Chen, Xi, et al.
Pubblicazione: (2026)
SARES-DEIM: Sparse Mixture-of-Experts Meets DETR for Robust SAR Ship Detection
di: Song, Fenghao, et al.
Pubblicazione: (2026)
di: Song, Fenghao, et al.
Pubblicazione: (2026)
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
di: Wang, Xin, et al.
Pubblicazione: (2026)
di: Wang, Xin, et al.
Pubblicazione: (2026)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence
di: She, Chaoyin, et al.
Pubblicazione: (2025)
di: She, Chaoyin, et al.
Pubblicazione: (2025)
Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
di: He, Landi, et al.
Pubblicazione: (2026)
di: He, Landi, et al.
Pubblicazione: (2026)
LadderMoE: Ladder-Side Mixture of Experts Adapters for Bronze Inscription Recognition
di: Zhou, Rixin, et al.
Pubblicazione: (2025)
di: Zhou, Rixin, et al.
Pubblicazione: (2025)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
di: Kim, Sohee, et al.
Pubblicazione: (2025)
di: Kim, Sohee, et al.
Pubblicazione: (2025)
Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts
di: Chen, Qizhou, et al.
Pubblicazione: (2024)
di: Chen, Qizhou, et al.
Pubblicazione: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
di: Zong, Zhuofan, et al.
Pubblicazione: (2024)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
di: Liang, Xiao, et al.
Pubblicazione: (2025)
di: Liang, Xiao, et al.
Pubblicazione: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
di: Cai, Kaitong, et al.
Pubblicazione: (2025)
di: Cai, Kaitong, et al.
Pubblicazione: (2025)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2024)
di: Zhang, Yue, et al.
Pubblicazione: (2024)
EVLM: An Efficient Vision-Language Model for Visual Understanding
di: Chen, Kaibing, et al.
Pubblicazione: (2024)
di: Chen, Kaibing, et al.
Pubblicazione: (2024)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
di: Zhang, Ce, et al.
Pubblicazione: (2025)
di: Zhang, Ce, et al.
Pubblicazione: (2025)
LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
di: Wang, Yimu, et al.
Pubblicazione: (2025)
di: Wang, Yimu, et al.
Pubblicazione: (2025)
Towards Vision Mixture of Experts for Wildlife Monitoring on the Edge
di: Mensah, Emmanuel Azuh, et al.
Pubblicazione: (2024)
di: Mensah, Emmanuel Azuh, et al.
Pubblicazione: (2024)
SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane Recognition
di: Cai, Qing, et al.
Pubblicazione: (2025)
di: Cai, Qing, et al.
Pubblicazione: (2025)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
A Survey of Token Compression for Efficient Multimodal Large Language Models
di: Shao, Kele, et al.
Pubblicazione: (2025)
di: Shao, Kele, et al.
Pubblicazione: (2025)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
di: Chen, Junjie, et al.
Pubblicazione: (2025)
di: Chen, Junjie, et al.
Pubblicazione: (2025)
Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
di: Gao, Qinghao, et al.
Pubblicazione: (2025)
di: Gao, Qinghao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
di: Cai, Chaoxiang, et al.
Pubblicazione: (2025) -
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
di: Jia, Chenwei, et al.
Pubblicazione: (2026) -
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024) -
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
di: Cao, Sihan, et al.
Pubblicazione: (2026) -
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
di: Wang, Dianyi, et al.
Pubblicazione: (2025)