MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Libo, He, Peixiong, Harn, Po-Wei, Qin, Xiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
Minimal-Intervention KV Retention via Set-Conditioned Diversity
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Horseshoe Mixtures-of-Experts (HS-MoE)
von: Polson, Nick, et al.
Veröffentlicht: (2026)
von: Polson, Nick, et al.
Veröffentlicht: (2026)
LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning
von: Rodriguez, Ariel, et al.
Veröffentlicht: (2026)
von: Rodriguez, Ariel, et al.
Veröffentlicht: (2026)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
MoE-Loco: Mixture of Experts for Multitask Locomotion
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
von: Shi, Long, et al.
Veröffentlicht: (2025)
von: Shi, Long, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant Computing
von: Xiao, Wenjing, et al.
Veröffentlicht: (2025)
von: Xiao, Wenjing, et al.
Veröffentlicht: (2025)
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
von: Jin, Haopeng
Veröffentlicht: (2026)
von: Jin, Haopeng
Veröffentlicht: (2026)
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
von: Jin, In-Hwan, et al.
Veröffentlicht: (2025)
von: Jin, In-Hwan, et al.
Veröffentlicht: (2025)
MoE3D: A Mixture-of-Experts Module for 3D Reconstruction
von: Wang, Zichen, et al.
Veröffentlicht: (2026)
von: Wang, Zichen, et al.
Veröffentlicht: (2026)
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
von: Salehi, Mohammad Reza Deylam, et al.
Veröffentlicht: (2026)
von: Salehi, Mohammad Reza Deylam, et al.
Veröffentlicht: (2026)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
von: Feng, Jiarui, et al.
Veröffentlicht: (2026)
von: Feng, Jiarui, et al.
Veröffentlicht: (2026)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
von: Sun, Libo, et al.
Veröffentlicht: (2026) -
Minimal-Intervention KV Retention via Set-Conditioned Diversity
von: Sun, Libo, et al.
Veröffentlicht: (2026) -
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024) -
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
von: Chen, Xiaodong, et al.
Veröffentlicht: (2025) -
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)