Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Zhiyuan, Hong, Zicong, Huang, Yuegui, Lyu, Yufeng, Chen, Wuhui, Yu, Yue, Yu, Fan, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
by: Fang, Zhiyuan, et al.
Published: (2025)
by: Fang, Zhiyuan, et al.
Published: (2025)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
by: Huang, Yuegui, et al.
Published: (2026)
by: Huang, Yuegui, et al.
Published: (2026)
Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
by: Wen, Junyi, et al.
Published: (2025)
by: Wen, Junyi, et al.
Published: (2025)
Training and Serving System of Foundation Models: A Comprehensive Survey
by: Zhou, Jiahang, et al.
Published: (2024)
by: Zhou, Jiahang, et al.
Published: (2024)
GriDB: Scaling Blockchain Database via Sharding and Off-Chain Cross-Shard Mechanism
by: Hong, Zicong, et al.
Published: (2024)
by: Hong, Zicong, et al.
Published: (2024)
DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
by: Bai, Sikai, et al.
Published: (2025)
by: Bai, Sikai, et al.
Published: (2025)
Mixture-of-Experts for Distributed Edge Computing with Channel-Aware Gating Function
by: Song, Qiuchen, et al.
Published: (2025)
by: Song, Qiuchen, et al.
Published: (2025)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025)
by: Zheng, Chen, et al.
Published: (2025)
On Bayesian Softmax-Gated Mixture-of-Experts Models
by: Bariletto, Nicola, et al.
Published: (2026)
by: Bariletto, Nicola, et al.
Published: (2026)
Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
by: Qin, Shengling, et al.
Published: (2025)
by: Qin, Shengling, et al.
Published: (2025)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
by: Huang, Minbin, et al.
Published: (2026)
by: Huang, Minbin, et al.
Published: (2026)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
by: Yan, Jiaming, et al.
Published: (2025)
by: Yan, Jiaming, et al.
Published: (2025)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
Gaussian Process-Gated Hierarchical Mixtures of Experts
by: Liu, Yuhao, et al.
Published: (2023)
by: Liu, Yuhao, et al.
Published: (2023)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
by: Jin, Zewen, et al.
Published: (2025)
by: Jin, Zewen, et al.
Published: (2025)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
by: Jiang, Fan, et al.
Published: (2026)
by: Jiang, Fan, et al.
Published: (2026)
Mixture-of-Experts Diffusion Models for Adaptive Massive MIMO Channel Estimation via Variational Bayesian Inference
by: Jiang, Zhuorui, et al.
Published: (2026)
by: Jiang, Zhuorui, et al.
Published: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
M$^4$oE: A Foundation Model for Medical Multimodal Image Segmentation with Mixture of Experts
by: Jiang, Yufeng, et al.
Published: (2024)
by: Jiang, Yufeng, et al.
Published: (2024)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
Federated Mixture-of-Expert for Non-Overlapped Cross-Domain Sequential Recommendation
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
The Osteoblastic Microenvironment Determines the Fate of Breast Cancer Cells Disseminated in the Bone Marrow
by: Hong‐Li Wang, et al.
Published: (2026)
by: Hong‐Li Wang, et al.
Published: (2026)
Similar Items
-
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
by: Fang, Zhiyuan, et al.
Published: (2025) -
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
by: Huang, Yuegui, et al.
Published: (2026) -
Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
by: Wen, Junyi, et al.
Published: (2025) -
Training and Serving System of Foundation Models: A Comprehensive Survey
by: Zhou, Jiahang, et al.
Published: (2024) -
GriDB: Scaling Blockchain Database via Sharding and Off-Chain Cross-Shard Mechanism
by: Hong, Zicong, et al.
Published: (2024)