Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Bowen, Shen, Yikang, Liu, Haokun, Mishra, Mayank, Zhang, Gaoyuan, Oliva, Aude, Raffel, Colin, Panda, Rameswar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
Scattered Mixture-of-Experts Implementation
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
Position: The Most Expensive Part of an LLM should be its Training Data
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
Soft Merging of Experts with Adaptive Routing
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
Training Sparse Mixture Of Experts Text Embedding Models
von: Nussbaum, Zach, et al.
Veröffentlicht: (2025)
von: Nussbaum, Zach, et al.
Veröffentlicht: (2025)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
von: Feng, Yuchen, et al.
Veröffentlicht: (2025)
von: Feng, Yuchen, et al.
Veröffentlicht: (2025)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Rethinking Sparse Mixture of Experts from a Unified Perspective
von: Do, Giang, et al.
Veröffentlicht: (2025)
von: Do, Giang, et al.
Veröffentlicht: (2025)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2024)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2024)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
von: Deng, Haikang, et al.
Veröffentlicht: (2023)
von: Deng, Haikang, et al.
Veröffentlicht: (2023)
Enhancing Training Data Attribution with Representational Optimization
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
von: Wang, Peiqi, et al.
Veröffentlicht: (2024)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Structured Code Representations Enable Data-Efficient Adaptation of Code Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2024)
von: Agarwal, Mayank, et al.
Veröffentlicht: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
von: Shen, Ke, et al.
Veröffentlicht: (2024)
von: Shen, Ke, et al.
Veröffentlicht: (2024)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
SEAP: Training-free Sparse Expert Activation Pruning Unlock the Brainpower of Large Language Models
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
von: Kang, Junmo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025) -
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023) -
Scattered Mixture-of-Experts Implementation
von: Tan, Shawn, et al.
Veröffentlicht: (2024) -
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024) -
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)