Routing-Free Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yilun, Han, Jinru, Yan, Sikuan, Tresp, Volker, Ma, Yunpu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
von: Liao, Ruotong, et al.
Veröffentlicht: (2023)
von: Liao, Ruotong, et al.
Veröffentlicht: (2023)
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
von: Ding, Zifeng, et al.
Veröffentlicht: (2023)
von: Ding, Zifeng, et al.
Veröffentlicht: (2023)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
von: Niu, Tianyi, et al.
Veröffentlicht: (2026)
von: Niu, Tianyi, et al.
Veröffentlicht: (2026)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
von: Bi, Jinhe, et al.
Veröffentlicht: (2025)
von: Bi, Jinhe, et al.
Veröffentlicht: (2025)
Bayes or Heisenberg: Who(se) Rules?
von: Tresp, Volker, et al.
Veröffentlicht: (2025)
von: Tresp, Volker, et al.
Veröffentlicht: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
von: Ma, Xiaowen, et al.
Veröffentlicht: (2025)
von: Ma, Xiaowen, et al.
Veröffentlicht: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
von: Liu, Feilong
Veröffentlicht: (2026)
von: Liu, Feilong
Veröffentlicht: (2026)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System
von: Liu, Mingyan
Veröffentlicht: (2025)
von: Liu, Mingyan
Veröffentlicht: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
von: Hallee, Logan, et al.
Veröffentlicht: (2024)
von: Hallee, Logan, et al.
Veröffentlicht: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025) -
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024) -
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
von: Liao, Ruotong, et al.
Veröffentlicht: (2023) -
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
von: Ding, Zifeng, et al.
Veröffentlicht: (2023) -
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)