$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Takashiro, Shota, Kojima, Takeshi, Taniguchi, Shohei, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
por: Gambardella, Andrew, et al.
Publicado: (2025)
por: Gambardella, Andrew, et al.
Publicado: (2025)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
por: Takashiro, Shota, et al.
Publicado: (2024)
por: Takashiro, Shota, et al.
Publicado: (2024)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
por: Cao, Qi, et al.
Publicado: (2026)
por: Cao, Qi, et al.
Publicado: (2026)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024)
por: Muzio, Alexandre, et al.
Publicado: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
por: Gu, Naibin, et al.
Publicado: (2025)
por: Gu, Naibin, et al.
Publicado: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
por: Gambardella, Andrew, et al.
Publicado: (2024)
por: Gambardella, Andrew, et al.
Publicado: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
por: Ludziejewski, Jan, et al.
Publicado: (2025)
por: Ludziejewski, Jan, et al.
Publicado: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
por: Chen, Guanjie, et al.
Publicado: (2024)
por: Chen, Guanjie, et al.
Publicado: (2024)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
por: Liu, Yang, et al.
Publicado: (2026)
por: Liu, Yang, et al.
Publicado: (2026)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
por: Sun, Weigao, et al.
Publicado: (2025)
por: Sun, Weigao, et al.
Publicado: (2025)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
por: Kang, Hao, et al.
Publicado: (2025)
por: Kang, Hao, et al.
Publicado: (2025)
Steering MoE LLMs via Expert (De)Activation
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
por: Li, Pingzhi, et al.
Publicado: (2025)
por: Li, Pingzhi, et al.
Publicado: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
por: Xia, Xinfeng, et al.
Publicado: (2025)
por: Xia, Xinfeng, et al.
Publicado: (2025)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
por: Sun, Libo, et al.
Publicado: (2026)
por: Sun, Libo, et al.
Publicado: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
por: Yuan, Yueming, et al.
Publicado: (2025)
por: Yuan, Yueming, et al.
Publicado: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
por: Deng, Jianing, et al.
Publicado: (2026)
por: Deng, Jianing, et al.
Publicado: (2026)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
por: Huang, Quzhe, et al.
Publicado: (2024)
por: Huang, Quzhe, et al.
Publicado: (2024)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
por: Li, Zichong, et al.
Publicado: (2025)
por: Li, Zichong, et al.
Publicado: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
por: Xie, Yanyue, et al.
Publicado: (2024)
por: Xie, Yanyue, et al.
Publicado: (2024)
ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns
por: Zhao, Ziyu, et al.
Publicado: (2026)
por: Zhao, Ziyu, et al.
Publicado: (2026)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
por: Kang, Junmo, et al.
Publicado: (2024)
por: Kang, Junmo, et al.
Publicado: (2024)
MoFE: Mixture of Frozen Experts Architecture
por: Seo, Jean, et al.
Publicado: (2025)
por: Seo, Jean, et al.
Publicado: (2025)
Dynamic Injection of Entity Knowledge into Dense Retrievers
por: Yamada, Ikuya, et al.
Publicado: (2025)
por: Yamada, Ikuya, et al.
Publicado: (2025)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
por: Harada, Keno, et al.
Publicado: (2025)
por: Harada, Keno, et al.
Publicado: (2025)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
por: Kojima, Takeshi, et al.
Publicado: (2024)
por: Kojima, Takeshi, et al.
Publicado: (2024)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
por: Shi, Chufan, et al.
Publicado: (2024)
por: Shi, Chufan, et al.
Publicado: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
por: Chen, Yanbei, et al.
Publicado: (2026)
por: Chen, Yanbei, et al.
Publicado: (2026)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
por: Ye, Charles, et al.
Publicado: (2026)
por: Ye, Charles, et al.
Publicado: (2026)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
por: Su, Zhenpeng, et al.
Publicado: (2024)
por: Su, Zhenpeng, et al.
Publicado: (2024)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
por: Zhou, Wenyong, et al.
Publicado: (2026)
por: Zhou, Wenyong, et al.
Publicado: (2026)
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
por: Xu, Yuzhuang, et al.
Publicado: (2025)
por: Xu, Yuzhuang, et al.
Publicado: (2025)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
por: Matsutani, Kohsei, et al.
Publicado: (2025)
por: Matsutani, Kohsei, et al.
Publicado: (2025)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
por: Lv, Xingtai, et al.
Publicado: (2026)
por: Lv, Xingtai, et al.
Publicado: (2026)
Ejemplares similares
-
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
por: Gambardella, Andrew, et al.
Publicado: (2025) -
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
por: Takashiro, Shota, et al.
Publicado: (2024) -
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
por: Cao, Qi, et al.
Publicado: (2026) -
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024) -
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)