PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jung, Min Jae, Kim, JooHee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Superposition in Transformers: A Novel Way of Building Mixture of Experts
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
Public Data Assisted Differentially Private In-Context Learning
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
von: Lyu, Boxuan, et al.
Veröffentlicht: (2026)
von: Lyu, Boxuan, et al.
Veröffentlicht: (2026)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
von: Pham, Loc, et al.
Veröffentlicht: (2026)
von: Pham, Loc, et al.
Veröffentlicht: (2026)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
A Two-Step Approach for Data-Efficient French Pronunciation Learning
von: Lee, Hoyeon, et al.
Veröffentlicht: (2024)
von: Lee, Hoyeon, et al.
Veröffentlicht: (2024)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning
von: Shi, Boyan, et al.
Veröffentlicht: (2026)
von: Shi, Boyan, et al.
Veröffentlicht: (2026)
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Routing-Free Mixture-of-Experts
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
von: Liu, Boan, et al.
Veröffentlicht: (2023)
von: Liu, Boan, et al.
Veröffentlicht: (2023)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
von: Hallee, Logan, et al.
Veröffentlicht: (2024)
von: Hallee, Logan, et al.
Veröffentlicht: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
von: Moon, Sehwan, et al.
Veröffentlicht: (2025)
von: Moon, Sehwan, et al.
Veröffentlicht: (2025)
Yuan 2.0-M32: Mixture of Experts with Attention Router
von: Wu, Shaohua, et al.
Veröffentlicht: (2024)
von: Wu, Shaohua, et al.
Veröffentlicht: (2024)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026) -
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025) -
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025) -
Superposition in Transformers: A Novel Way of Building Mixture of Experts
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024) -
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)