Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Anrui, Huang, Ruijun, Zhang, Xin, Dong, Fang, Cao, Hengjie, Huang, Zhendong, Yang, Yifeng, Chen, Mengyi, Zhou, Jixian, Dong, Mingzhi, Wang, Yujiang, Hou, Jinlong, Lv, Qin, Dick, Robert P., Cheng, Yuan, Lu, Tun, Yang, Fan, Shang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SD-MoE: Spectral Decomposition for Effective Expert Specialization
by: Huang, Ruijun, et al.
Published: (2026)
by: Huang, Ruijun, et al.
Published: (2026)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026)
by: Huang, Zhendong, et al.
Published: (2026)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
Dispelling the Curse of Singularities in Neural Network Optimizations
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
MoE-CT: A Novel Approach For Large Language Models Training With Resistance To Catastrophic Forgetting
by: Li, Tianhao, et al.
Published: (2024)
by: Li, Tianhao, et al.
Published: (2024)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
by: Li, Junzhuo, et al.
Published: (2025)
by: Li, Junzhuo, et al.
Published: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
by: Shi, Yubin, et al.
Published: (2024)
by: Shi, Yubin, et al.
Published: (2024)
Catastrophic Forgetting in Kolmogorov-Arnold Networks
by: Rahman, Mohammad Marufur, et al.
Published: (2025)
by: Rahman, Mohammad Marufur, et al.
Published: (2025)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
MuseumMaker: Continual Style Customization without Catastrophic Forgetting
by: Liu, Chenxi, et al.
Published: (2024)
by: Liu, Chenxi, et al.
Published: (2024)
Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
by: Xiong, Yifeng, et al.
Published: (2025)
by: Xiong, Yifeng, et al.
Published: (2025)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
by: Wu, Haoyuan, et al.
Published: (2025)
by: Wu, Haoyuan, et al.
Published: (2025)
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
by: Huang, Jianheng, et al.
Published: (2024)
by: Huang, Jianheng, et al.
Published: (2024)
SiftMoE: Similarity-Aware Energy-Efficient Expert Selection for Wireless Distributed MoE Inference
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Avoid Catastrophic Forgetting with Rank-1 Fisher from Diffusion Models
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information Density
by: Mi, Zhendong, et al.
Published: (2026)
by: Mi, Zhendong, et al.
Published: (2026)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
by: Huang, Quzhe, et al.
Published: (2024)
by: Huang, Quzhe, et al.
Published: (2024)
Continual Learning and Catastrophic Forgetting
by: van de Ven, Gido M., et al.
Published: (2024)
by: van de Ven, Gido M., et al.
Published: (2024)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025)
by: Peng, Ze, et al.
Published: (2025)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
by: Zhang, Zhendong
Published: (2025)
by: Zhang, Zhendong
Published: (2025)
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting
by: Chen, Haolin, et al.
Published: (2024)
by: Chen, Haolin, et al.
Published: (2024)
Overcoming Catastrophic Forgetting by Exemplar Selection in Task-oriented Dialogue System
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026)
by: Cui, Chenwei, et al.
Published: (2026)
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
by: Koubbi, Hugo, et al.
Published: (2024)
by: Koubbi, Hugo, et al.
Published: (2024)
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
by: Sun, Xingwu, et al.
Published: (2024)
by: Sun, Xingwu, et al.
Published: (2024)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
MReg: A Novel Regression Model with MoE-based Video Feature Mining for Mitral Regurgitation Diagnosis
by: Liu, Zhe, et al.
Published: (2025)
by: Liu, Zhe, et al.
Published: (2025)
Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
by: Xu, Xinrun, et al.
Published: (2025)
by: Xu, Xinrun, et al.
Published: (2025)
Conditions for Catastrophic Forgetting in Multilingual Translation
by: Liu, Danni, et al.
Published: (2025)
by: Liu, Danni, et al.
Published: (2025)
Predicting the Susceptibility of Examples to Catastrophic Forgetting
by: Hacohen, Guy, et al.
Published: (2024)
by: Hacohen, Guy, et al.
Published: (2024)
Similar Items
-
SD-MoE: Spectral Decomposition for Effective Expert Specialization
by: Huang, Ruijun, et al.
Published: (2026) -
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026) -
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026) -
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025) -
Dispelling the Curse of Singularities in Neural Network Optimizations
by: Cao, Hengjie, et al.
Published: (2026)