MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tastan, Nurbek, Laskaridis, Stefanos, Nandakumar, Karthik, Horvath, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aequa: Fair Model Rewards in Collaborative Learning via Slimmable Networks
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
Stochastic Self-Organization in Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
CYCle: Choosing Your Collaborators Wisely to Enhance Collaborative Fairness in Decentralized Learning
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
FedPeWS: Personalized Warmup via Subnetworks for Enhanced Heterogeneous Federated Learning
by: Tastan, Nurbek, et al.
Published: (2024)
by: Tastan, Nurbek, et al.
Published: (2024)
A Framework for Double-Blind Federated Adaptation of Foundation Models
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026)
by: Tastan, Nurbek, et al.
Published: (2026)
MOLM: Mixture of LoRA Markers
by: Fares, Samar, et al.
Published: (2025)
by: Fares, Samar, et al.
Published: (2025)
SPDMark: Selective Parameter Displacement for Robust Video Watermarking
by: Fares, Samar, et al.
Published: (2025)
by: Fares, Samar, et al.
Published: (2025)
Redefining Contributions: Shapley-Driven Federated Learning
by: Tastan, Nurbek, et al.
Published: (2024)
by: Tastan, Nurbek, et al.
Published: (2024)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
by: Zaccone, Riccardo, et al.
Published: (2026)
by: Zaccone, Riccardo, et al.
Published: (2026)
Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning
by: Ukaye, Asim, et al.
Published: (2026)
by: Ukaye, Asim, et al.
Published: (2026)
MoSE: Unveiling Structural Patterns in Graphs via Mixture of Subgraph Experts
by: Ye, Junda, et al.
Published: (2025)
by: Ye, Junda, et al.
Published: (2025)
Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New Baseline
by: Al-lahham, Anas, et al.
Published: (2024)
by: Al-lahham, Anas, et al.
Published: (2024)
Data-Free Contribution Estimation in Federated Learning using Gradient von Neumann Entropy
by: Ukaye, Asim, et al.
Published: (2026)
by: Ukaye, Asim, et al.
Published: (2026)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
by: Liu, Zehua, et al.
Published: (2025)
by: Liu, Zehua, et al.
Published: (2025)
Maestro: Uncovering Low-Rank Structures via Trainable Decomposition
by: Horvath, Samuel, et al.
Published: (2023)
by: Horvath, Samuel, et al.
Published: (2023)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
MoFE: Mixture of Frozen Experts Architecture
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
by: Chen, Guanjie, et al.
Published: (2024)
by: Chen, Guanjie, et al.
Published: (2024)
Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation
by: Wei, Jia, et al.
Published: (2025)
by: Wei, Jia, et al.
Published: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
by: Feng, Jinyuan, et al.
Published: (2025)
by: Feng, Jinyuan, et al.
Published: (2025)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
by: Muzio, Alexandre, et al.
Published: (2024)
by: Muzio, Alexandre, et al.
Published: (2024)
Bayesian Mixture of Experts For Large Language Models
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
by: Zeng, Hanqing, et al.
Published: (2025)
by: Zeng, Hanqing, et al.
Published: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
by: Gao, Shangqian, et al.
Published: (2025)
by: Gao, Shangqian, et al.
Published: (2025)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
by: Ai, Mengting, et al.
Published: (2023)
by: Ai, Mengting, et al.
Published: (2023)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
A Survey on Mixture of Experts in Large Language Models
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
by: Wang, An, et al.
Published: (2024)
by: Wang, An, et al.
Published: (2024)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
by: Yoon, Youngsik, et al.
Published: (2026)
by: Yoon, Youngsik, et al.
Published: (2026)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Beyond instruction-conditioning, MoTE: Mixture of Task Experts for Multi-task Embedding Models
by: Romero, Miguel, et al.
Published: (2025)
by: Romero, Miguel, et al.
Published: (2025)
A Closer Look into Mixture-of-Experts in Large Language Models
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Similar Items
-
Aequa: Fair Model Rewards in Collaborative Learning via Slimmable Networks
by: Tastan, Nurbek, et al.
Published: (2025) -
Stochastic Self-Organization in Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2025) -
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
by: Tastan, Nurbek, et al.
Published: (2025) -
CYCle: Choosing Your Collaborators Wisely to Enhance Collaborative Fairness in Decentralized Learning
by: Tastan, Nurbek, et al.
Published: (2025) -
FedPeWS: Personalized Warmup via Subnetworks for Enhanced Heterogeneous Federated Learning
by: Tastan, Nurbek, et al.
Published: (2024)