When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Libo, Harn, Po-wei, He, Peixiong, Qin, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
by: Sun, Libo, et al.
Published: (2026)
by: Sun, Libo, et al.
Published: (2026)
Minimal-Intervention KV Retention via Set-Conditioned Diversity
by: Sun, Libo, et al.
Published: (2026)
by: Sun, Libo, et al.
Published: (2026)
Nucleus-Image: Sparse MoE for Image Generation
by: Akiti, Chandan, et al.
Published: (2026)
by: Akiti, Chandan, et al.
Published: (2026)
Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
by: Bao, Wei, et al.
Published: (2026)
by: Bao, Wei, et al.
Published: (2026)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
by: Min, Chengxi, et al.
Published: (2025)
by: Min, Chengxi, et al.
Published: (2025)
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
by: Kada, Masahiro, et al.
Published: (2026)
by: Kada, Masahiro, et al.
Published: (2026)
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
by: Zheng, Youwei, et al.
Published: (2025)
by: Zheng, Youwei, et al.
Published: (2025)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
by: Qin, Guangshuo, et al.
Published: (2026)
by: Qin, Guangshuo, et al.
Published: (2026)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026)
by: Jeddi, Ahmadreza, et al.
Published: (2026)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
by: Ma, Yingjie, et al.
Published: (2024)
by: Ma, Yingjie, et al.
Published: (2024)
MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration
by: Kong, Lingshun, et al.
Published: (2026)
by: Kong, Lingshun, et al.
Published: (2026)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Sparse Spectral LoRA: Routed Experts for Medical VLMs
by: Manzari, Omid Nejati, et al.
Published: (2026)
by: Manzari, Omid Nejati, et al.
Published: (2026)
GNN-MoE: Context-Aware Patch Routing using GNNs for Parameter-Efficient Domain Generalization
by: Soliman, Mahmoud, et al.
Published: (2025)
by: Soliman, Mahmoud, et al.
Published: (2025)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
HiFi-MambaV2: Hierarchical Shared-Routed MoE for High-Fidelity MRI Reconstruction
by: Fang, Pengcheng, et al.
Published: (2025)
by: Fang, Pengcheng, et al.
Published: (2025)
When are Diffusion Priors Helpful in Sparse Reconstruction? A Study with Sparse-view CT
by: Cheung, Matt Y., et al.
Published: (2025)
by: Cheung, Matt Y., et al.
Published: (2025)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2026)
by: Ko, Dohwan, et al.
Published: (2026)
SocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-Tuning
by: Kawabata, Tomohito, et al.
Published: (2025)
by: Kawabata, Tomohito, et al.
Published: (2025)
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
by: Xiang, Qi, et al.
Published: (2025)
by: Xiang, Qi, et al.
Published: (2025)
AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
SparseLaneSTP: Leveraging Spatio-Temporal Priors with Sparse Transformers for 3D Lane Detection
by: Pittner, Maximilian, et al.
Published: (2026)
by: Pittner, Maximilian, et al.
Published: (2026)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Vision Transformer with Sparse Scan Prior
by: Zhang, Yuguang, et al.
Published: (2024)
by: Zhang, Yuguang, et al.
Published: (2024)
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
by: Tang, Pin, et al.
Published: (2024)
by: Tang, Pin, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
by: Jin, In-Hwan, et al.
Published: (2025)
by: Jin, In-Hwan, et al.
Published: (2025)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
by: Zhu, Xingkui, et al.
Published: (2024)
by: Zhu, Xingkui, et al.
Published: (2024)
When Does Pruning Benefit Vision Representations?
by: Cassano, Enrico, et al.
Published: (2025)
by: Cassano, Enrico, et al.
Published: (2025)
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
by: Shi, Yangming, et al.
Published: (2026)
by: Shi, Yangming, et al.
Published: (2026)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
by: Zhang, Qizhe, et al.
Published: (2023)
by: Zhang, Qizhe, et al.
Published: (2023)
Leveraging Sparse Annotations for Leukemia Diagnosis on the Large Leukemia Dataset
by: Rehman, Abdul, et al.
Published: (2025)
by: Rehman, Abdul, et al.
Published: (2025)
OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning
by: Geng, Xinyu, et al.
Published: (2024)
by: Geng, Xinyu, et al.
Published: (2024)
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
by: Qin, Chuan, et al.
Published: (2026)
by: Qin, Chuan, et al.
Published: (2026)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
Similar Items
-
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
by: Sun, Libo, et al.
Published: (2026) -
Minimal-Intervention KV Retention via Set-Conditioned Diversity
by: Sun, Libo, et al.
Published: (2026) -
Nucleus-Image: Sparse MoE for Image Generation
by: Akiti, Chandan, et al.
Published: (2026) -
Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
by: Bao, Wei, et al.
Published: (2026) -
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
by: Wei, Yujie, et al.
Published: (2025)