MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Novikov, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
von: Tang, Anke, et al.
Veröffentlicht: (2024)
von: Tang, Anke, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Zero-shot Generalizable Graph Anomaly Detection with Mixture of Riemannian Experts
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
MoEMeta: Mixture-of-Experts Meta Learning for Few-Shot Relational Learning
von: Wu, Han, et al.
Veröffentlicht: (2025)
von: Wu, Han, et al.
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
MicroNAS: Zero-Shot Neural Architecture Search for MCUs
von: Qiao, Ye, et al.
Veröffentlicht: (2024)
von: Qiao, Ye, et al.
Veröffentlicht: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
Zero-Shot Robustification of Zero-Shot Models
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Constructing Efficient Fact-Storing MLPs for Transformers
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2025)
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2025)
Neural Metamorphosis
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Sparsity and Superposition in Mixture of Experts
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
Mixture of Diverse Size Experts
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
Mixture of Concept Bottleneck Experts
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
Mixture of A Million Experts
von: He, Xu Owen
Veröffentlicht: (2024)
von: He, Xu Owen
Veröffentlicht: (2024)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
Weight-based Decomposition: A Case for Bilinear MLPs
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Rank Also Matters: Hierarchical Configuration for Mixture of Adapter Experts in LLM Fine-Tuning
von: Cong, Peizhuang, et al.
Veröffentlicht: (2025)
von: Cong, Peizhuang, et al.
Veröffentlicht: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback
von: Jedidi, Nour, et al.
Veröffentlicht: (2024)
von: Jedidi, Nour, et al.
Veröffentlicht: (2024)
Mixture of Experts in Large Language Models
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
Graph Knowledge Distillation to Mixture of Experts
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
Mixture of Weak & Strong Experts on Graphs
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024) -
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025) -
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
von: Tang, Anke, et al.
Veröffentlicht: (2024) -
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026) -
Zero-shot Generalizable Graph Anomaly Detection with Mixture of Riemannian Experts
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)