MoVA: Adapting Mixture of Vision Experts to Multimodal Context
Fuente:
arXiv
Saved in:
| Main Authors: | Zong, Zhuofan, Ma, Bingqi, Shen, Dazhong, Song, Guanglu, Shao, Hao, Jiang, Dongzhi, Li, Hongsheng, Liu, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
by: Zong, Zhuofan, et al.
Published: (2024)
by: Zong, Zhuofan, et al.
Published: (2024)
ADT: Tuning Diffusion Models with Adversarial Supervision
by: Shen, Dazhong, et al.
Published: (2025)
by: Shen, Dazhong, et al.
Published: (2025)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
by: Ma, Bingqi, et al.
Published: (2024)
by: Ma, Bingqi, et al.
Published: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023)
by: Xue, Zeyue, et al.
Published: (2023)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
by: Shen, Leyang, et al.
Published: (2024)
by: Shen, Leyang, et al.
Published: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
by: He, Dailan, et al.
Published: (2026)
by: He, Dailan, et al.
Published: (2026)
Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation
by: Wang, Fu-Yun, et al.
Published: (2024)
by: Wang, Fu-Yun, et al.
Published: (2024)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
by: Shen, Dazhong, et al.
Published: (2024)
by: Shen, Dazhong, et al.
Published: (2024)
Adapted-MoE: Mixture of Experts with Test-Time Adaption for Anomaly Detection
by: Lei, Tianwu, et al.
Published: (2024)
by: Lei, Tianwu, et al.
Published: (2024)
MoME: Mixture of Multimodal Experts for Cancer Survival Prediction
by: Xiong, Conghao, et al.
Published: (2024)
by: Xiong, Conghao, et al.
Published: (2024)
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
by: Shao, Hao, et al.
Published: (2026)
by: Shao, Hao, et al.
Published: (2026)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
by: Chen, Shaoxiang, et al.
Published: (2024)
by: Chen, Shaoxiang, et al.
Published: (2024)
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
by: Liao, Minwen, et al.
Published: (2025)
by: Liao, Minwen, et al.
Published: (2025)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts
by: Han, Xumeng, et al.
Published: (2024)
by: Han, Xumeng, et al.
Published: (2024)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
by: Xue, Rongkun, et al.
Published: (2024)
by: Xue, Rongkun, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
by: Jiang, Ruixiang, et al.
Published: (2024)
by: Jiang, Ruixiang, et al.
Published: (2024)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding
by: Chopra, Shivang, et al.
Published: (2025)
by: Chopra, Shivang, et al.
Published: (2025)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
by: Lin, Hui, et al.
Published: (2024)
by: Lin, Hui, et al.
Published: (2024)
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
by: Ma, Yueen, et al.
Published: (2025)
by: Ma, Yueen, et al.
Published: (2025)
MoCTEFuse: Illumination-Gated Mixture of Chiral Transformer Experts for Multi-Level Infrared and Visible Image Fusion
by: Jinfu, Li, et al.
Published: (2025)
by: Jinfu, Li, et al.
Published: (2025)
Similar Items
-
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
by: Zong, Zhuofan, et al.
Published: (2024) -
ADT: Tuning Diffusion Models with Adversarial Supervision
by: Shen, Dazhong, et al.
Published: (2025) -
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
by: Ma, Bingqi, et al.
Published: (2024) -
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024) -
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
by: Shao, Hao, et al.
Published: (2024)