Dense Backpropagation Improves Training for Sparse Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Panda, Ashwinee, Baherwani, Vatsal, Sarwar, Zain, Therien, Benjamin, Sahu, Sambit, Goldstein, Tom, Chakraborty, Supriyo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912687099740160
author Panda, Ashwinee
Baherwani, Vatsal
Sarwar, Zain
Therien, Benjamin
Sahu, Sambit
Goldstein, Tom
Chakraborty, Supriyo
author_facet Panda, Ashwinee
Baherwani, Vatsal
Sarwar, Zain
Therien, Benjamin
Sahu, Sambit
Goldstein, Tom
Chakraborty, Supriyo
contents Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performance. We present a lightweight approximation method that gives the MoE router a dense gradient update while continuing to sparsely activate its parameters. Our method, which we refer to as Default MoE, substitutes missing expert activations with default outputs consisting of an exponential moving average of expert outputs previously seen over the course of training. This allows the router to receive signals from every expert for each token, leading to significant improvements in training performance. Our Default MoE outperforms standard TopK routing in a variety of settings without requiring significant computational overhead. Code: https://github.com/vatsal0/default-moe.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
Panda, Ashwinee
Baherwani, Vatsal
Sarwar, Zain
Therien, Benjamin
Sahu, Sambit
Goldstein, Tom
Chakraborty, Supriyo
Machine Learning
Artificial Intelligence
Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performance. We present a lightweight approximation method that gives the MoE router a dense gradient update while continuing to sparsely activate its parameters. Our method, which we refer to as Default MoE, substitutes missing expert activations with default outputs consisting of an exponential moving average of expert outputs previously seen over the course of training. This allows the router to receive signals from every expert for each token, leading to significant improvements in training performance. Our Default MoE outperforms standard TopK routing in a variety of settings without requiring significant computational overhead. Code: https://github.com/vatsal0/default-moe.
title Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.12463