The Empirical Impact of Reducing Symmetries on the Performance of Deep Ensembles and MoE
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chernov, Andrei, Novitskij, Oleg |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
(GG) MoE vs. MLP on Tabular Data
von: Chernov, Andrei
Veröffentlicht: (2025)
von: Chernov, Andrei
Veröffentlicht: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks
von: Chernov, Andrei
Veröffentlicht: (2025)
von: Chernov, Andrei
Veröffentlicht: (2025)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
Deep Ensembles Secretly Perform Empirical Bayes
von: Loaiza-Ganem, Gabriel, et al.
Veröffentlicht: (2025)
von: Loaiza-Ganem, Gabriel, et al.
Veröffentlicht: (2025)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
von: Duanmu, Haojie, et al.
Veröffentlicht: (2025)
von: Duanmu, Haojie, et al.
Veröffentlicht: (2025)
Spectral Manifold Regularization for Stable and Modular Routing in Deep MoE Architectures
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2026)
von: Xu, Zukang, et al.
Veröffentlicht: (2026)
DOT-MoE: Differentiable Optimal Transport for MoEfication
von: Bamba, Udbhav, et al.
Veröffentlicht: (2026)
von: Bamba, Udbhav, et al.
Veröffentlicht: (2026)
Fine-Tuning a Time Series Foundation Model with Wasserstein Loss
von: Chernov, Andrei
Veröffentlicht: (2024)
von: Chernov, Andrei
Veröffentlicht: (2024)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
von: Hu, Gang, et al.
Veröffentlicht: (2025)
von: Hu, Gang, et al.
Veröffentlicht: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
GRIN: GRadient-INformed MoE
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
MoEITS: A Green AI approach for simplifying MoE-LLMs
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2026)
von: Lau, Tim Tsz-Kit, et al.
Veröffentlicht: (2026)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
von: Su, Yang, et al.
Veröffentlicht: (2025)
von: Su, Yang, et al.
Veröffentlicht: (2025)
Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
von: Liao, Ning, et al.
Veröffentlicht: (2025)
von: Liao, Ning, et al.
Veröffentlicht: (2025)
STAMImputer: Spatio-Temporal Attention MoE for Traffic Data Imputation
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
von: Lim, Derek, et al.
Veröffentlicht: (2024)
von: Lim, Derek, et al.
Veröffentlicht: (2024)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
von: Tran, Huyen Ngoc, et al.
Veröffentlicht: (2026)
von: Tran, Huyen Ngoc, et al.
Veröffentlicht: (2026)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
(GG) MoE vs. MLP on Tabular Data
von: Chernov, Andrei
Veröffentlicht: (2025) -
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025) -
Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks
von: Chernov, Andrei
Veröffentlicht: (2025) -
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026) -
SDG-MoE: Signed Debate Graph Mixture-of-Experts
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)