Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Lou, Meng, Fu, Yunxiang, Yu, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation
by: Fu, Yunxiang, et al.
Published: (2024)
by: Fu, Yunxiang, et al.
Published: (2024)
SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
by: Lou, Meng, et al.
Published: (2024)
by: Lou, Meng, et al.
Published: (2024)
A2Mamba: Attention-augmented State Space Models for Visual Recognition
by: Lou, Meng, et al.
Published: (2025)
by: Lou, Meng, et al.
Published: (2025)
Stable Routing for Mixture-of-Experts in Class-Incremental Learning
by: Guo, Zirui, et al.
Published: (2026)
by: Guo, Zirui, et al.
Published: (2026)
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
by: Yuan, Yike, et al.
Published: (2025)
by: Yuan, Yike, et al.
Published: (2025)
EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts
by: Cheng, Anzhe, et al.
Published: (2026)
by: Cheng, Anzhe, et al.
Published: (2026)
Mixture-of-Experts Models in Vision: Routing, Optimization, and Generalization
by: Rokah, Adam, et al.
Published: (2026)
by: Rokah, Adam, et al.
Published: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
by: Jeon, Jaebyeong, et al.
Published: (2025)
by: Jeon, Jaebyeong, et al.
Published: (2025)
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Domain-Specialized Object Detection via Model-Level Mixtures of Experts
by: Pavlitska, Svetlana, et al.
Published: (2026)
by: Pavlitska, Svetlana, et al.
Published: (2026)
Adaptive Shared Experts with LoRA-Based Mixture of Experts for Multi-Task Learning
by: Yang, Minghao, et al.
Published: (2025)
by: Yang, Minghao, et al.
Published: (2025)
MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
by: Qiu, Zihuan, et al.
Published: (2025)
by: Qiu, Zihuan, et al.
Published: (2025)
Robust Emotion Recognition via Bi-Level Self-Supervised Continual Learning
by: Ahmad, Adnan, et al.
Published: (2025)
by: Ahmad, Adnan, et al.
Published: (2025)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging
by: Shen, Li, et al.
Published: (2024)
by: Shen, Li, et al.
Published: (2024)
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
Towards Adversarial Robustness of Model-Level Mixture-of-Experts Architectures for Semantic Segmentation
by: Pavlitska, Svetlana, et al.
Published: (2024)
by: Pavlitska, Svetlana, et al.
Published: (2024)
Task-conditioned Ensemble of Expert Models for Continuous Learning
by: Sharma, Renu, et al.
Published: (2025)
by: Sharma, Renu, et al.
Published: (2025)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
by: Oldfield, James, et al.
Published: (2024)
by: Oldfield, James, et al.
Published: (2024)
Mixture of Group Experts for Learning Invariant Representations
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Video Relationship Detection Using Mixture of Experts
by: Shaabana, Ala, et al.
Published: (2024)
by: Shaabana, Ala, et al.
Published: (2024)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
Task-Aware Multi-Expert Architecture For Lifelong Deep Learning
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
Mixture of Experts in Image Classification: What's the Sweet Spot?
by: Videau, Mathurin, et al.
Published: (2024)
by: Videau, Mathurin, et al.
Published: (2024)
FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual Learning
by: Yan, Hongwei, et al.
Published: (2026)
by: Yan, Hongwei, et al.
Published: (2026)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
by: Huai, Tianyu, et al.
Published: (2025)
by: Huai, Tianyu, et al.
Published: (2025)
Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
Bi-Level Unsupervised Feature Selection
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba
by: Fu, Yunxiang, et al.
Published: (2024)
by: Fu, Yunxiang, et al.
Published: (2024)
Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning
by: Wang, Huiyi, et al.
Published: (2024)
by: Wang, Huiyi, et al.
Published: (2024)
Doubly Perturbed Task Free Continual Learning
by: Lee, Byung Hyun, et al.
Published: (2023)
by: Lee, Byung Hyun, et al.
Published: (2023)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
by: Kowsher, Md, et al.
Published: (2026)
by: Kowsher, Md, et al.
Published: (2026)
Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
by: Liu, Yahui, et al.
Published: (2025)
by: Liu, Yahui, et al.
Published: (2025)
Lightweight Metadata-Aware Mixture-of-Experts Masked Autoencoder for Earth Observation
by: Albughdadi, Mohanad
Published: (2025)
by: Albughdadi, Mohanad
Published: (2025)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
by: Luo, Jun, et al.
Published: (2024)
by: Luo, Jun, et al.
Published: (2024)
Similar Items
-
Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning
by: Lou, Meng, et al.
Published: (2026) -
SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation
by: Fu, Yunxiang, et al.
Published: (2024) -
SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
by: Lou, Meng, et al.
Published: (2024) -
A2Mamba: Attention-augmented State Space Models for Visual Recognition
by: Lou, Meng, et al.
Published: (2025) -
Stable Routing for Mixture-of-Experts in Class-Incremental Learning
by: Guo, Zirui, et al.
Published: (2026)