CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Jun-Tao, Xie, Zhen-Hao, Shi, Yu-Cheng, Zhou, Da-Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
von: Tang, Jun-Tao, et al.
Veröffentlicht: (2026)
von: Tang, Jun-Tao, et al.
Veröffentlicht: (2026)
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
von: Shi, Yu-Cheng, et al.
Veröffentlicht: (2026)
von: Shi, Yu-Cheng, et al.
Veröffentlicht: (2026)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
von: Liang, Yunlong, et al.
Veröffentlicht: (2025)
von: Liang, Yunlong, et al.
Veröffentlicht: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
von: Zheng, Kening, et al.
Veröffentlicht: (2026)
von: Zheng, Kening, et al.
Veröffentlicht: (2026)
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
von: Pan, Lehan, et al.
Veröffentlicht: (2026)
von: Pan, Lehan, et al.
Veröffentlicht: (2026)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities
von: Xia, Guoyang, et al.
Veröffentlicht: (2025)
von: Xia, Guoyang, et al.
Veröffentlicht: (2025)
GMoE: Empowering LLMs Fine-Tuning via MoE Graph Collaboration
von: Bai, Ting, et al.
Veröffentlicht: (2024)
von: Bai, Ting, et al.
Veröffentlicht: (2024)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
ShardMemo: Masked MoE Routing for Sharded Agentic LLM Memory
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
von: Hu, Tao, et al.
Veröffentlicht: (2026)
von: Hu, Tao, et al.
Veröffentlicht: (2026)
Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2026)
Expert-Token Resonance MoE: Bidirectional Routing with Efficiency Affinity-Driven Active Selection
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Xinfeng, et al.
Veröffentlicht: (2025)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
von: Li, Zongqian, et al.
Veröffentlicht: (2025)
von: Li, Zongqian, et al.
Veröffentlicht: (2025)
Training Report of TeleChat3-MoE
von: Liu, Xinzhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinzhang, et al.
Veröffentlicht: (2025)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
Towards Automatic Continual Learning: A Self-Adaptive Framework for Continual Instruction Tuning
von: Lin, Peiyi, et al.
Veröffentlicht: (2025)
von: Lin, Peiyi, et al.
Veröffentlicht: (2025)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
GRIN: GRadient-INformed MoE
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration
von: Wei, Linye, et al.
Veröffentlicht: (2026)
von: Wei, Linye, et al.
Veröffentlicht: (2026)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
von: Tang, Jun-Tao, et al.
Veröffentlicht: (2026) -
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
von: Shi, Yu-Cheng, et al.
Veröffentlicht: (2026) -
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026) -
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025) -
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)