ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Chao, Wei, Xinming, Zhong, Yinmin, Yang, Chengxu, Wu, Bingyang, Zhu, Ruidong, Zhang, Zili, Liu, Yuliang, Jin, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Heddle: A Distributed Orchestration System for Agentic RL Rollout
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
von: Jin, Haopeng
Veröffentlicht: (2026)
von: Jin, Haopeng
Veröffentlicht: (2026)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
von: Dong, Daize, et al.
Veröffentlicht: (2026)
von: Dong, Daize, et al.
Veröffentlicht: (2026)
Optimizing RLHF Training for Large Language Models with Stage Fusion
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
Fine-grained MoE Load Balancing with Linear Programming
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
ShardMemo: Masked MoE Routing for Sharded Agentic LLM Memory
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
von: Shi, Long, et al.
Veröffentlicht: (2025)
von: Shi, Long, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
EMO: Frustratingly Easy Progressive Training of Extendable MoE
von: Jin, Linghao, et al.
Veröffentlicht: (2026)
von: Jin, Linghao, et al.
Veröffentlicht: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
von: Liang, Yunlong, et al.
Veröffentlicht: (2025)
von: Liang, Yunlong, et al.
Veröffentlicht: (2025)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
von: Zhu, Xingkui, et al.
Veröffentlicht: (2024)
von: Zhu, Xingkui, et al.
Veröffentlicht: (2024)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Heddle: A Distributed Orchestration System for Agentic RL Rollout
von: Zhang, Zili, et al.
Veröffentlicht: (2026) -
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
von: Jin, Haopeng
Veröffentlicht: (2026) -
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
von: Wu, Bingyang, et al.
Veröffentlicht: (2025) -
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
von: Dong, Daize, et al.
Veröffentlicht: (2026) -
Optimizing RLHF Training for Large Language Models with Stage Fusion
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)