Towards Faster Language Model Inference Using Mixture-of-Experts Flow Matching
Fuente:
arXiv
Salvato in:
| Autore principale: | Li, Aihua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026)
di: Madan, Vivan, et al.
Pubblicazione: (2026)
Mixture of Experts in Large Language Models
di: Zhang, Danyang, et al.
Pubblicazione: (2025)
di: Zhang, Danyang, et al.
Pubblicazione: (2025)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
di: Yan, Jiaming, et al.
Pubblicazione: (2025)
di: Yan, Jiaming, et al.
Pubblicazione: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
di: Gao, Yuting, et al.
Pubblicazione: (2025)
di: Gao, Yuting, et al.
Pubblicazione: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
di: Kim, Gyeongman, et al.
Pubblicazione: (2025)
di: Kim, Gyeongman, et al.
Pubblicazione: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
di: Pan, Bowen, et al.
Pubblicazione: (2024)
di: Pan, Bowen, et al.
Pubblicazione: (2024)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
di: Wang, Yun, et al.
Pubblicazione: (2025)
di: Wang, Yun, et al.
Pubblicazione: (2025)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
di: Raje, Arian, et al.
Pubblicazione: (2026)
di: Raje, Arian, et al.
Pubblicazione: (2026)
Unveiling Hidden Collaboration within Mixture-of-Experts in Large Language Models
di: Tang, Yuanbo, et al.
Pubblicazione: (2025)
di: Tang, Yuanbo, et al.
Pubblicazione: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
di: Jin, Zewen, et al.
Pubblicazione: (2025)
di: Jin, Zewen, et al.
Pubblicazione: (2025)
Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts
di: Miao, Changhao, et al.
Pubblicazione: (2026)
di: Miao, Changhao, et al.
Pubblicazione: (2026)
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fang, Zhiyuan, et al.
Pubblicazione: (2025)
Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion
di: Tang, Anke, et al.
Pubblicazione: (2024)
di: Tang, Anke, et al.
Pubblicazione: (2024)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
di: Park, Sumin, et al.
Pubblicazione: (2025)
di: Park, Sumin, et al.
Pubblicazione: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Mixture of Latent Experts Using Tensor Products
di: Su, Zhan, et al.
Pubblicazione: (2024)
di: Su, Zhan, et al.
Pubblicazione: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
di: He, Yifei, et al.
Pubblicazione: (2025)
di: He, Yifei, et al.
Pubblicazione: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
di: Chen, Yuanteng, et al.
Pubblicazione: (2025)
di: Chen, Yuanteng, et al.
Pubblicazione: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
di: Lu, Xudong, et al.
Pubblicazione: (2024)
di: Lu, Xudong, et al.
Pubblicazione: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
di: He, Neil, et al.
Pubblicazione: (2025)
di: He, Neil, et al.
Pubblicazione: (2025)
Upcycling Large Language Models into Mixture of Experts
di: He, Ethan, et al.
Pubblicazione: (2024)
di: He, Ethan, et al.
Pubblicazione: (2024)
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
di: Liang, Jingcong, et al.
Pubblicazione: (2025)
di: Liang, Jingcong, et al.
Pubblicazione: (2025)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
di: He, Shwai, et al.
Pubblicazione: (2025)
di: He, Shwai, et al.
Pubblicazione: (2025)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
di: Liu, Baihui, et al.
Pubblicazione: (2026)
di: Liu, Baihui, et al.
Pubblicazione: (2026)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
Mixture of Raytraced Experts
di: Perin, Andrea, et al.
Pubblicazione: (2025)
di: Perin, Andrea, et al.
Pubblicazione: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
Tokenised Flow Matching for Hierarchical Simulation Based Inference
di: Charles, Giovanni, et al.
Pubblicazione: (2026)
di: Charles, Giovanni, et al.
Pubblicazione: (2026)
OLMoE: Open Mixture-of-Experts Language Models
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
di: Yao, Jinghan, et al.
Pubblicazione: (2024)
di: Yao, Jinghan, et al.
Pubblicazione: (2024)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
di: Vankov, Daniil, et al.
Pubblicazione: (2026)
di: Vankov, Daniil, et al.
Pubblicazione: (2026)
Mixture of Experts in a Mixture of RL settings
di: Willi, Timon, et al.
Pubblicazione: (2024)
di: Willi, Timon, et al.
Pubblicazione: (2024)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
di: Li, Junzhuo, et al.
Pubblicazione: (2026)
di: Li, Junzhuo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026) -
Mixture of Experts in Large Language Models
di: Zhang, Danyang, et al.
Pubblicazione: (2025) -
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
di: Yan, Jiaming, et al.
Pubblicazione: (2025) -
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025) -
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
di: Gao, Yuting, et al.
Pubblicazione: (2025)