FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Hao, Yu, Zichun, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
PithTrain: A Compact and Agent-Native MoE Training System
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
von: Yu, Zichun, et al.
Veröffentlicht: (2024)
von: Yu, Zichun, et al.
Veröffentlicht: (2024)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
von: Muzio, Alexandre, et al.
Veröffentlicht: (2024)
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
von: Li, Xiaochuan, et al.
Veröffentlicht: (2024)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2024)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
von: Li, Zichong, et al.
Veröffentlicht: (2025)
von: Li, Zichong, et al.
Veröffentlicht: (2025)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
von: Zhao, Ziyu, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2026)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2026)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
Group-Level Data Selection for Efficient Pretraining
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
End-to-End Ontology Learning with Large Language Models
von: Lo, Andy, et al.
Veröffentlicht: (2024)
von: Lo, Andy, et al.
Veröffentlicht: (2024)
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2025)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2025)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
von: Su, Zhenpeng, et al.
Veröffentlicht: (2024)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2024)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
von: Yu, Zichun, et al.
Veröffentlicht: (2025) -
PithTrain: A Compact and Agent-Native MoE Training System
von: Lai, Ruihang, et al.
Veröffentlicht: (2026) -
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
von: Yu, Zichun, et al.
Veröffentlicht: (2024) -
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
von: Yu, Zichun, et al.
Veröffentlicht: (2026) -
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)