Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinbo, Jia, Shian, Huang, Ziyang, Cao, Jing, Song, Mingli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent Centric Operating System -- a Comprehensive Review and Outlook for Operating System
von: Jia, Shian, et al.
Veröffentlicht: (2024)
von: Jia, Shian, et al.
Veröffentlicht: (2024)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
Duration-Informed Workload Scheduler
von: Loreti, Daniela, et al.
Veröffentlicht: (2026)
von: Loreti, Daniela, et al.
Veröffentlicht: (2026)
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
von: Hu, Yi-Xiang, et al.
Veröffentlicht: (2026)
von: Hu, Yi-Xiang, et al.
Veröffentlicht: (2026)
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
von: Wang, Xiaoye
Veröffentlicht: (2024)
von: Wang, Xiaoye
Veröffentlicht: (2024)
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
Workload Schedulers -- Genesis, Algorithms and Differences
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
TRAIL: Trust-Aware Client Scheduling for Semi-Decentralized Federated Learning
von: Hu, Gangqiang, et al.
Veröffentlicht: (2024)
von: Hu, Gangqiang, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review
von: Gu, Yan, et al.
Veröffentlicht: (2025)
von: Gu, Yan, et al.
Veröffentlicht: (2025)
CoRaiS: Lightweight Real-Time Scheduler for Multi-Edge Cooperative Computing
von: Hu, Yujiao, et al.
Veröffentlicht: (2024)
von: Hu, Yujiao, et al.
Veröffentlicht: (2024)
Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems
von: Sgambati, Matthew, et al.
Veröffentlicht: (2025)
von: Sgambati, Matthew, et al.
Veröffentlicht: (2025)
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration
von: Patra, Sunandita, et al.
Veröffentlicht: (2025)
von: Patra, Sunandita, et al.
Veröffentlicht: (2025)
TS-EoH: An Edge Server Task Scheduling Algorithm Based on Evolution of Heuristic
von: Yatong, Wang, et al.
Veröffentlicht: (2024)
von: Yatong, Wang, et al.
Veröffentlicht: (2024)
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
WORKSWORLD: A Domain for Integrated Numeric Planning and Scheduling of Distributed Pipelined Workflows
von: Paul, Taylor, et al.
Veröffentlicht: (2026)
von: Paul, Taylor, et al.
Veröffentlicht: (2026)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Application of Machine Learning Optimization in Cloud Computing Resource Scheduling and Management
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center Scheduling
von: Pathak, Nilavra, et al.
Veröffentlicht: (2026)
von: Pathak, Nilavra, et al.
Veröffentlicht: (2026)
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
von: Li, Boyang, et al.
Veröffentlicht: (2024)
von: Li, Boyang, et al.
Veröffentlicht: (2024)
Online Client Scheduling and Resource Allocation for Efficient Federated Edge Learning
von: Gao, Zhidong, et al.
Veröffentlicht: (2024)
von: Gao, Zhidong, et al.
Veröffentlicht: (2024)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Agent Centric Operating System -- a Comprehensive Review and Outlook for Operating System
von: Jia, Shian, et al.
Veröffentlicht: (2024) -
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025) -
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025) -
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)