Superposition in Transformers: A Novel Way of Building Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Chaliah, Ayoub Ben, Dellagi, Hela |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
by: Chaliah, Ayoub Ben, et al.
Published: (2025)
by: Chaliah, Ayoub Ben, et al.
Published: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025)
by: Bershatsky, Daniel, et al.
Published: (2025)
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
by: Petridis, Savvas, et al.
Published: (2024)
by: Petridis, Savvas, et al.
Published: (2024)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
by: Jung, Min Jae, et al.
Published: (2024)
by: Jung, Min Jae, et al.
Published: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
by: Lyu, Boxuan, et al.
Published: (2026)
by: Lyu, Boxuan, et al.
Published: (2026)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
by: Zhu, Shien, et al.
Published: (2025)
by: Zhu, Shien, et al.
Published: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
by: Roy, Sampurna, et al.
Published: (2025)
by: Roy, Sampurna, et al.
Published: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
by: Sun, Liangtai, et al.
Published: (2024)
by: Sun, Liangtai, et al.
Published: (2024)
Multi-Head Mixture-of-Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
by: Liu, Boan, et al.
Published: (2023)
by: Liu, Boan, et al.
Published: (2023)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
by: Lee, Wonjun, et al.
Published: (2026)
by: Lee, Wonjun, et al.
Published: (2026)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
by: Pham, Loc, et al.
Published: (2026)
by: Pham, Loc, et al.
Published: (2026)
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts
by: Du, Wenzhuo, et al.
Published: (2025)
by: Du, Wenzhuo, et al.
Published: (2025)
Sparsity and Superposition in Mixture of Experts
by: Chaudhari, Marmik, et al.
Published: (2025)
by: Chaudhari, Marmik, et al.
Published: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
Yuan 2.0-M32: Mixture of Experts with Attention Router
by: Wu, Shaohua, et al.
Published: (2024)
by: Wu, Shaohua, et al.
Published: (2024)
Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics
by: Hou, Iyad Ait, et al.
Published: (2026)
by: Hou, Iyad Ait, et al.
Published: (2026)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Investigating the potential of Sparse Mixtures-of-Experts for multi-domain neural machine translation
by: Chirkova, Nadezhda, et al.
Published: (2024)
by: Chirkova, Nadezhda, et al.
Published: (2024)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
by: Zhuang, Yuan, et al.
Published: (2025)
by: Zhuang, Yuan, et al.
Published: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
by: Deng, Guanzhi, et al.
Published: (2026)
by: Deng, Guanzhi, et al.
Published: (2026)
Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
by: Ternovtsii, Ivan, et al.
Published: (2025)
by: Ternovtsii, Ivan, et al.
Published: (2025)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
by: Nguyen, Nam V., et al.
Published: (2025)
by: Nguyen, Nam V., et al.
Published: (2025)
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
by: NVIDIA, et al.
Published: (2026)
by: NVIDIA, et al.
Published: (2026)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
by: Wei, Tianwen, et al.
Published: (2024)
by: Wei, Tianwen, et al.
Published: (2024)
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
by: Cui, Christopher Z., et al.
Published: (2024)
by: Cui, Christopher Z., et al.
Published: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Similar Items
-
Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
by: Chaliah, Ayoub Ben, et al.
Published: (2025) -
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025) -
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
by: Petridis, Savvas, et al.
Published: (2024) -
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
by: Jung, Min Jae, et al.
Published: (2024) -
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)