PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Min Jae, Kim, JooHee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
by: Herbst, Jeremy, et al.
Published: (2026)
by: Herbst, Jeremy, et al.
Published: (2026)
Mixtures of SubExperts for Large Language Continual Learning
by: Kang, Haeyong
Published: (2025)
by: Kang, Haeyong
Published: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025)
by: Bershatsky, Daniel, et al.
Published: (2025)
Superposition in Transformers: A Novel Way of Building Mixture of Experts
by: Chaliah, Ayoub Ben, et al.
Published: (2024)
by: Chaliah, Ayoub Ben, et al.
Published: (2024)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
by: Lee, Wonjun, et al.
Published: (2026)
by: Lee, Wonjun, et al.
Published: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
Public Data Assisted Differentially Private In-Context Learning
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
OLMoE: Open Mixture-of-Experts Language Models
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
by: Lyu, Boxuan, et al.
Published: (2026)
by: Lyu, Boxuan, et al.
Published: (2026)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
by: Kim, Gyeongman, et al.
Published: (2025)
by: Kim, Gyeongman, et al.
Published: (2025)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
by: Pham, Loc, et al.
Published: (2026)
by: Pham, Loc, et al.
Published: (2026)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
by: Nguyen, Nam V., et al.
Published: (2025)
by: Nguyen, Nam V., et al.
Published: (2025)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
by: Zhu, Shien, et al.
Published: (2025)
by: Zhu, Shien, et al.
Published: (2025)
A Two-Step Approach for Data-Efficient French Pronunciation Learning
by: Lee, Hoyeon, et al.
Published: (2024)
by: Lee, Hoyeon, et al.
Published: (2024)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
by: Jiang, Fan, et al.
Published: (2026)
by: Jiang, Fan, et al.
Published: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
by: Chen, Yanbei, et al.
Published: (2026)
by: Chen, Yanbei, et al.
Published: (2026)
SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning
by: Shi, Boyan, et al.
Published: (2026)
by: Shi, Boyan, et al.
Published: (2026)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
by: Roy, Sampurna, et al.
Published: (2025)
by: Roy, Sampurna, et al.
Published: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Multi-Head Mixture-of-Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
by: Wei, Tianwen, et al.
Published: (2024)
by: Wei, Tianwen, et al.
Published: (2024)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
by: Petridis, Savvas, et al.
Published: (2024)
by: Petridis, Savvas, et al.
Published: (2024)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
by: Liu, Boan, et al.
Published: (2023)
by: Liu, Boan, et al.
Published: (2023)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
by: Hallee, Logan, et al.
Published: (2024)
by: Hallee, Logan, et al.
Published: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)
by: Ban, Minjeong, et al.
Published: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
Yuan 2.0-M32: Mixture of Experts with Attention Router
by: Wu, Shaohua, et al.
Published: (2024)
by: Wu, Shaohua, et al.
Published: (2024)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
by: Sun, Liangtai, et al.
Published: (2024)
by: Sun, Liangtai, et al.
Published: (2024)
Similar Items
-
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
by: Herbst, Jeremy, et al.
Published: (2026) -
Mixtures of SubExperts for Large Language Continual Learning
by: Kang, Haeyong
Published: (2025) -
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025) -
Superposition in Transformers: A Novel Way of Building Mixture of Experts
by: Chaliah, Ayoub Ben, et al.
Published: (2024) -
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
by: Liu, Yang, et al.
Published: (2026)