Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hao Mark, Mo, Zhiwen, Lee, Royson, Wang, Qianzhou, Li, Da, Hu, Shell Xu, Luk, Wayne, Hospedales, Timothy, Fan, Hongxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Model Diffusion for Certifiable Few-shot Transfer Learning
von: Rezk, Fady, et al.
Veröffentlicht: (2025)
von: Rezk, Fady, et al.
Veröffentlicht: (2025)
Feed-Forward Latent Domain Adaptation
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2022)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2022)
MobileQuant: Mobile-friendly Quantization for On-device Language Models
von: Tan, Fuwen, et al.
Veröffentlicht: (2024)
von: Tan, Fuwen, et al.
Veröffentlicht: (2024)
FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Recurrent Early Exits for Federated Learning with Heterogeneous Clients
von: Lee, Royson, et al.
Veröffentlicht: (2024)
von: Lee, Royson, et al.
Veröffentlicht: (2024)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
von: Lee, Royson, et al.
Veröffentlicht: (2025)
von: Lee, Royson, et al.
Veröffentlicht: (2025)
A Bayesian Approach to Data Point Selection
von: Xu, Xinnuo, et al.
Veröffentlicht: (2024)
von: Xu, Xinnuo, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error Correction
von: Campbell, Charlie, et al.
Veröffentlicht: (2025)
von: Campbell, Charlie, et al.
Veröffentlicht: (2025)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
von: Zhang, Zehuan, et al.
Veröffentlicht: (2024)
von: Zhang, Zehuan, et al.
Veröffentlicht: (2024)
ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
von: Chavhan, Ruchika, et al.
Veröffentlicht: (2024)
von: Chavhan, Ruchika, et al.
Veröffentlicht: (2024)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
TradExpert: Revolutionizing Trading with Mixture of Expert LLMs
von: Ding, Qianggang, et al.
Veröffentlicht: (2024)
von: Ding, Qianggang, et al.
Veröffentlicht: (2024)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting
von: Liu, Fan, et al.
Veröffentlicht: (2025)
von: Liu, Fan, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Mixture of Experts for Low-Resource LLMs
von: Joseph, Ori Bar, et al.
Veröffentlicht: (2026)
von: Joseph, Ori Bar, et al.
Veröffentlicht: (2026)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
von: Hao, Jiawei, et al.
Veröffentlicht: (2026)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
von: Bai, Jun, et al.
Veröffentlicht: (2025)
von: Bai, Jun, et al.
Veröffentlicht: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Adaptive Shared Experts with LoRA-Based Mixture of Experts for Multi-Task Learning
von: Yang, Minghao, et al.
Veröffentlicht: (2025)
von: Yang, Minghao, et al.
Veröffentlicht: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Accelerating 3D Gaussian Splatting with Neural Sorting and Axis-Oriented Rasterization
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025) -
Model Diffusion for Certifiable Few-shot Transfer Learning
von: Rezk, Fady, et al.
Veröffentlicht: (2025) -
Feed-Forward Latent Domain Adaptation
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2022) -
MobileQuant: Mobile-friendly Quantization for On-device Language Models
von: Tan, Fuwen, et al.
Veröffentlicht: (2024) -
FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)