Robustness of Mixtures of Experts to Feature Noise
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Dong, Nittala, Rahul, Burkholz, Rebekka |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?
by: Bause, Franka, et al.
Published: (2026)
by: Bause, Franka, et al.
Published: (2026)
Fixed Aggregation Features Can Rival GNNs
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024)
by: Jacobs, Tom, et al.
Published: (2024)
GATE: How to Keep Out Intrusive Neighbors
by: Mustafa, Nimrah, et al.
Published: (2024)
by: Mustafa, Nimrah, et al.
Published: (2024)
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
GNNs Getting ComFy: Community and Feature Similarity Guided Rewiring
by: Rubio-Madrigal, Celia, et al.
Published: (2025)
by: Rubio-Madrigal, Celia, et al.
Published: (2025)
HORST: Composing Optimizer Geometries for Sparse Transformer Training
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Spectral Graph Pruning Against Over-Squashing and Over-Smoothing
by: Jamadandi, Adarsh, et al.
Published: (2024)
by: Jamadandi, Adarsh, et al.
Published: (2024)
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Pay Attention to Small Weights
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
Pruning neural network models for gene regulatory dynamics using data and domain knowledge
by: Hossain, Intekhab, et al.
Published: (2024)
by: Hossain, Intekhab, et al.
Published: (2024)
Frequency-Based Hyperparameter Selection in Games
by: Sanyal, Aniket, et al.
Published: (2026)
by: Sanyal, Aniket, et al.
Published: (2026)
When Shift Happens - Confounding Is to Blame
by: Reddy, Abbavaram Gowtham, et al.
Published: (2025)
by: Reddy, Abbavaram Gowtham, et al.
Published: (2025)
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
by: Adnan, Mohammed, et al.
Published: (2026)
by: Adnan, Mohammed, et al.
Published: (2026)
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Bridging Domains through Subspace-Aware Model Merging
by: Chaves, Levy, et al.
Published: (2026)
by: Chaves, Levy, et al.
Published: (2026)
Maximum Score Routing For Mixture-of-Experts
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Prediction-powered Inference by Mixture of Experts
by: Gu, Yanwu, et al.
Published: (2026)
by: Gu, Yanwu, et al.
Published: (2026)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
Mixture of Diverse Size Experts
by: Sun, Manxi, et al.
Published: (2024)
by: Sun, Manxi, et al.
Published: (2024)
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
by: Sun, Yuan
Published: (2025)
by: Sun, Yuan
Published: (2025)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
by: Muzio, Alexandre, et al.
Published: (2024)
by: Muzio, Alexandre, et al.
Published: (2024)
Path-Constrained Mixture-of-Experts
by: Gu, Zijin, et al.
Published: (2026)
by: Gu, Zijin, et al.
Published: (2026)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Impact of Label Noise on Learning Complex Features
by: Vashisht, Rahul, et al.
Published: (2024)
by: Vashisht, Rahul, et al.
Published: (2024)
RoME: Domain-Robust Mixture-of-Experts for MILP Solution Prediction across Domains
by: Pu, Tianle, et al.
Published: (2025)
by: Pu, Tianle, et al.
Published: (2025)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Mixture-of-Experts Meets In-Context Reinforcement Learning
by: Wu, Wenhao, et al.
Published: (2025)
by: Wu, Wenhao, et al.
Published: (2025)
Mixture of Raytraced Experts
by: Perin, Andrea, et al.
Published: (2025)
by: Perin, Andrea, et al.
Published: (2025)
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
DEGNN: Dual Experts Graph Neural Network Handling Both Edge and Node Feature Noise
by: Hasegawa, Tai, et al.
Published: (2024)
by: Hasegawa, Tai, et al.
Published: (2024)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Similar Items
-
Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?
by: Bause, Franka, et al.
Published: (2026) -
Fixed Aggregation Features Can Rival GNNs
by: Rubio-Madrigal, Celia, et al.
Published: (2026) -
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024) -
GATE: How to Keep Out Intrusive Neighbors
by: Mustafa, Nimrah, et al.
Published: (2024) -
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024)