Gespeichert in:
| 1. Verfasser: | Kim, Donghu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.18987 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
von: Han, Isaac, et al.
Veröffentlicht: (2026)
von: Han, Isaac, et al.
Veröffentlicht: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
von: Zou, Will Y., et al.
Veröffentlicht: (2025)
von: Zou, Will Y., et al.
Veröffentlicht: (2025)
Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
von: Lee, Hojoon, et al.
Veröffentlicht: (2024)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
von: Li, Junzhuo, et al.
Veröffentlicht: (2026)
von: Li, Junzhuo, et al.
Veröffentlicht: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Sparsity and Superposition in Mixture of Experts
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
Mixture of Diverse Size Experts
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
Mixture of Concept Bottleneck Experts
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
Mixture of A Million Experts
von: He, Xu Owen
Veröffentlicht: (2024)
von: He, Xu Owen
Veröffentlicht: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
von: Kim, Donghu, et al.
Veröffentlicht: (2024)
von: Kim, Donghu, et al.
Veröffentlicht: (2024)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
Adaptive Individual Uncertainty under Out-Of-Distribution Shift with Expert-Routed Conformal Prediction
von: Badkul, Amitesh, et al.
Veröffentlicht: (2025)
von: Badkul, Amitesh, et al.
Veröffentlicht: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
Mixture of Experts in Large Language Models
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
Graph Knowledge Distillation to Mixture of Experts
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
Mixture of Weak & Strong Experts on Graphs
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
Theory on Mixture-of-Experts in Continual Learning
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
von: Xue, Nan, et al.
Veröffentlicht: (2024)
von: Xue, Nan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026) -
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025) -
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
von: Kim, Hyunseung, et al.
Veröffentlicht: (2024) -
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025) -
FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
von: Han, Isaac, et al.
Veröffentlicht: (2026)