MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Pióro, Maciej, Ciebiera, Kamil, Król, Krystian, Ludziejewski, Jan, Krutul, Michał, Krajewski, Jakub, Antoniak, Szymon, Miłoś, Piotr, Cygan, Marek, Jaszczur, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023)
by: Antoniak, Szymon, et al.
Published: (2023)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
by: Krajewski, Jakub, et al.
Published: (2025)
by: Krajewski, Jakub, et al.
Published: (2025)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Patch-MoE Mamba: A Patch-Ordered Mixture-of-Experts State Space Architecture for Medical Image Segmentation
by: Adame, Diego, et al.
Published: (2026)
by: Adame, Diego, et al.
Published: (2026)
On the Theory of Risk-Aware Agents: Bridging Actor-Critic and Economics
by: Nauman, Michal, et al.
Published: (2023)
by: Nauman, Michal, et al.
Published: (2023)
Magnushammer: A Transformer-Based Approach to Premise Selection
by: Mikuła, Maciej, et al.
Published: (2023)
by: Mikuła, Maciej, et al.
Published: (2023)
Horseshoe Mixtures-of-Experts (HS-MoE)
by: Polson, Nick, et al.
Published: (2026)
by: Polson, Nick, et al.
Published: (2026)
Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation
by: Płotka, Szymon, et al.
Published: (2025)
by: Płotka, Szymon, et al.
Published: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
RoboMorph: Evolving Robot Morphology using Large Language Models
by: Qiu, Kevin, et al.
Published: (2024)
by: Qiu, Kevin, et al.
Published: (2024)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
by: Chen, Yuanteng, et al.
Published: (2025)
by: Chen, Yuanteng, et al.
Published: (2025)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
by: Huang, Runhan, et al.
Published: (2025)
by: Huang, Runhan, et al.
Published: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
by: Chen, Xiaodong, et al.
Published: (2025)
by: Chen, Xiaodong, et al.
Published: (2025)
Structured Packing in LLM Training Improves Long Context Utilization
by: Staniszewski, Konrad, et al.
Published: (2023)
by: Staniszewski, Konrad, et al.
Published: (2023)
Conversations on Mind, Matter, and Mathematics. Jean-Pierre Changeux and Alain Connes. Edited and translated by M. B. DeBevoise. Princeton, New Jersey: Princeton University Press, 1995, 261 p.; 15 x 22 cm. Glossary + Index. Language: English. ISBN: 0-691-08759-8.
by: Marek Antoniak
Published: (2013)
by: Marek Antoniak
Published: (2013)
Mixture of Experts (MoE): A Big Data Perspective
by: Gan, Wensheng, et al.
Published: (2025)
by: Gan, Wensheng, et al.
Published: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
by: Kulibaba, Stepan, et al.
Published: (2026)
by: Kulibaba, Stepan, et al.
Published: (2026)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
by: Xu, Yuhao, et al.
Published: (2026)
by: Xu, Yuhao, et al.
Published: (2026)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
by: Jin, In-Hwan, et al.
Published: (2025)
by: Jin, In-Hwan, et al.
Published: (2025)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
by: Feng, Jiarui, et al.
Published: (2026)
by: Feng, Jiarui, et al.
Published: (2026)
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
by: Zawalski, Michał, et al.
Published: (2024)
by: Zawalski, Michał, et al.
Published: (2024)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
by: Aghdam, Maryam Akhavan, et al.
Published: (2024)
by: Aghdam, Maryam Akhavan, et al.
Published: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
by: Muzio, Alexandre, et al.
Published: (2024)
by: Muzio, Alexandre, et al.
Published: (2024)
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025)
by: Balmau, Oana, et al.
Published: (2025)
Analyzing Internal Activity and Robustness of SNNs Across Neuron Parameter Space
by: Mazurek, Szymon, et al.
Published: (2025)
by: Mazurek, Szymon, et al.
Published: (2025)
Solutions at vacuum and rarefaction waves in pressureless Euler alignment system
by: Cygan, Szymon, et al.
Published: (2024)
by: Cygan, Szymon, et al.
Published: (2024)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
A Case for Validation Buffer in Pessimistic Actor-Critic
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
by: Su, Yang, et al.
Published: (2025)
by: Su, Yang, et al.
Published: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Astro-MoE: Mixture of Experts for Multiband Astronomical Time Series
by: Cádiz-Leyton, Martina, et al.
Published: (2025)
by: Cádiz-Leyton, Martina, et al.
Published: (2025)
Similar Items
-
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
by: Antoniak, Szymon, et al.
Published: (2023) -
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024) -
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025) -
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025) -
Decoupled Relative Learning Rate Schedules
by: Ludziejewski, Jan, et al.
Published: (2025)