Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Panda, Ashwinee, Baherwani, Vatsal, Sarwar, Zain, Therien, Benjamin, Sahu, Sambit, Goldstein, Tom, Chakraborty, Supriyo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025)
by: Zhang, Juzheng, et al.
Published: (2025)
Critique-Guided Distillation for Robust Reasoning via Refinement
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
by: Mircea, Andrei, et al.
Published: (2025)
by: Mircea, Andrei, et al.
Published: (2025)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
by: Nikolic, Strahinja, et al.
Published: (2025)
by: Nikolic, Strahinja, et al.
Published: (2025)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
by: Niu, Tianyi, et al.
Published: (2026)
by: Niu, Tianyi, et al.
Published: (2026)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
On Understanding of the Dynamics of Model Capacity in Continual Learning
by: Chakraborty, Supriyo, et al.
Published: (2025)
by: Chakraborty, Supriyo, et al.
Published: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
by: Rüb, Marcus, et al.
Published: (2024)
by: Rüb, Marcus, et al.
Published: (2024)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
by: Rastegar, Reza
Published: (2026)
by: Rastegar, Reza
Published: (2026)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
by: Zou, Will Y., et al.
Published: (2025)
by: Zou, Will Y., et al.
Published: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
by: Chamma, Ahmad, et al.
Published: (2025)
by: Chamma, Ahmad, et al.
Published: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
by: Novikov, Ivan
Published: (2025)
by: Novikov, Ivan
Published: (2025)
Dynamic Spectral Backpropagation for Efficient Neural Network Training
by: Muthuraman, Mannmohan
Published: (2025)
by: Muthuraman, Mannmohan
Published: (2025)
From Sparse to Soft Mixtures of Experts
by: Puigcerver, Joan, et al.
Published: (2023)
by: Puigcerver, Joan, et al.
Published: (2023)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
by: Swaroop, Anand, et al.
Published: (2025)
by: Swaroop, Anand, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
SAL: Selective Adaptive Learning for Backpropagation-Free Training with Sparsification
by: Liu, Fanping, et al.
Published: (2026)
by: Liu, Fanping, et al.
Published: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Mixture of Raytraced Experts
by: Perin, Andrea, et al.
Published: (2025)
by: Perin, Andrea, et al.
Published: (2025)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
RobotKeyframing: Learning Locomotion with High-Level Objectives via Mixture of Dense and Sparse Rewards
by: Zargarbashi, Fatemeh, et al.
Published: (2024)
by: Zargarbashi, Fatemeh, et al.
Published: (2024)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
by: Kunwar, Pradip, et al.
Published: (2025)
by: Kunwar, Pradip, et al.
Published: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Similar Items
-
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026) -
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025) -
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024) -
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025) -
Critique-Guided Distillation for Robust Reasoning via Refinement
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)