Multi-Layer Scheduling for MoE-Based LLM Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Sun, Yifan, Haffari, Gholamreza, Xu, Minxian, Buyya, Rajkumar, Toosi, Adel N. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
por: Bai, Xu, et al.
Publicado: (2025)
por: Bai, Xu, et al.
Publicado: (2025)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
por: Bai, Haoyu, et al.
Publicado: (2026)
por: Bai, Haoyu, et al.
Publicado: (2026)
DistributedEstimator: Distributed Training of Quantum Neural Networks via Circuit Cutting
por: Singh, Prabhjot, et al.
Publicado: (2026)
por: Singh, Prabhjot, et al.
Publicado: (2026)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
por: Wen, Linfeng, et al.
Publicado: (2024)
por: Wen, Linfeng, et al.
Publicado: (2024)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
por: Bai, Haoyu, et al.
Publicado: (2024)
por: Bai, Haoyu, et al.
Publicado: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
por: Liao, Junhan, et al.
Publicado: (2025)
por: Liao, Junhan, et al.
Publicado: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
por: Zhang, Zhexiang, et al.
Publicado: (2025)
por: Zhang, Zhexiang, et al.
Publicado: (2025)
A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments
por: Jayanetti, Amanda, et al.
Publicado: (2024)
por: Jayanetti, Amanda, et al.
Publicado: (2024)
Reinforcement Learning based Workflow Scheduling in Cloud and Edge Computing Environments: A Taxonomy, Review and Future Directions
por: Jayanetti, Amanda, et al.
Publicado: (2024)
por: Jayanetti, Amanda, et al.
Publicado: (2024)
A Knowledge Distillation-empowered Adaptive Federated Reinforcement Learning Framework for Multi-Domain IoT Applications Scheduling
por: Wang, Zhiyu, et al.
Publicado: (2025)
por: Wang, Zhiyu, et al.
Publicado: (2025)
TF-DDRL: A Transformer-enhanced Distributed DRL Technique for Scheduling IoT Applications in Edge and Cloud Computing Environments
por: Wang, Zhiyu, et al.
Publicado: (2024)
por: Wang, Zhiyu, et al.
Publicado: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
A Multi-Armed Bandit-Based Participant Selection Method for Federated Recommendation Systems
por: Liu, Jintao, et al.
Publicado: (2025)
por: Liu, Jintao, et al.
Publicado: (2025)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
por: Su, Zijie, et al.
Publicado: (2026)
por: Su, Zijie, et al.
Publicado: (2026)
A Joint Time and Energy-Efficient Federated Learning-based Computation Offloading Method for Mobile Edge Computing
por: Mukherjee, Anwesha, et al.
Publicado: (2024)
por: Mukherjee, Anwesha, et al.
Publicado: (2024)
EnFed: An Energy-aware Federated Learning in Resource Constrained Environments for Human Activity Recognition
por: Mukherjee, Anwesha, et al.
Publicado: (2024)
por: Mukherjee, Anwesha, et al.
Publicado: (2024)
Generative Federated Learning for Smart Prediction and Recommendation Applications
por: Mukherjee, Anwesha, et al.
Publicado: (2025)
por: Mukherjee, Anwesha, et al.
Publicado: (2025)
TrustMesh: A Blockchain-Enabled Trusted Distributed Computing Framework for Open Heterogeneous IoT Environments
por: Rangwala, Murtaza, et al.
Publicado: (2024)
por: Rangwala, Murtaza, et al.
Publicado: (2024)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
por: Yu, Shibo, et al.
Publicado: (2025)
por: Yu, Shibo, et al.
Publicado: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
por: Zhao, Lu, et al.
Publicado: (2025)
por: Zhao, Lu, et al.
Publicado: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)
por: Qian, Yulei, et al.
Publicado: (2024)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
por: Zhou, Guangyao, et al.
Publicado: (2021)
por: Zhou, Guangyao, et al.
Publicado: (2021)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
por: Wang, Haodong, et al.
Publicado: (2025)
por: Wang, Haodong, et al.
Publicado: (2025)
iDynamics: A Configurable Emulation Framework for Evaluating Microservice Scheduling Policies under Controllable Cloud-Edge Dynamics
por: Chen, Ming, et al.
Publicado: (2025)
por: Chen, Ming, et al.
Publicado: (2025)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
por: Akbari, Negin, et al.
Publicado: (2025)
por: Akbari, Negin, et al.
Publicado: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
por: Lee, Gunjun, et al.
Publicado: (2025)
por: Lee, Gunjun, et al.
Publicado: (2025)
A Risk-Aware UAV-Edge Service Framework for Wildfire Monitoring and Emergency Response
por: Huang, Yulun, et al.
Publicado: (2026)
por: Huang, Yulun, et al.
Publicado: (2026)
ReinFog: A Deep Reinforcement Learning Empowered Framework for Resource Management in Edge and Cloud Computing Environments
por: Wang, Zhiyu, et al.
Publicado: (2024)
por: Wang, Zhiyu, et al.
Publicado: (2024)
Placement of Microservices-based IoT Applications in Fog Computing: A Taxonomy and Future Directions
por: Pallewatta, Samodha, et al.
Publicado: (2022)
por: Pallewatta, Samodha, et al.
Publicado: (2022)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
por: Tang, Lujie, et al.
Publicado: (2024)
por: Tang, Lujie, et al.
Publicado: (2024)
Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference
por: Hewage, Tharindu B., et al.
Publicado: (2025)
por: Hewage, Tharindu B., et al.
Publicado: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
por: Nie, Xiaonan, et al.
Publicado: (2024)
por: Nie, Xiaonan, et al.
Publicado: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
por: Chen, Liangkun, et al.
Publicado: (2025)
por: Chen, Liangkun, et al.
Publicado: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
por: Doucet, Zachary, et al.
Publicado: (2025)
por: Doucet, Zachary, et al.
Publicado: (2025)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
por: Chen, Zihao, et al.
Publicado: (2025)
por: Chen, Zihao, et al.
Publicado: (2025)
A Cascaded Graph Neural Network for Joint Root Cause Localization and Analysis in Edge Computing Environments
por: Fernando, Duneesha, et al.
Publicado: (2026)
por: Fernando, Duneesha, et al.
Publicado: (2026)
A Decentralized Root Cause Localization Approach for Edge Computing Environments
por: Fernando, Duneesha, et al.
Publicado: (2025)
por: Fernando, Duneesha, et al.
Publicado: (2025)
A Hybrid Reactive-Proactive Auto-scaling Algorithm for SLA-Constrained Edge Computing
por: Gupta, Suhrid, et al.
Publicado: (2025)
por: Gupta, Suhrid, et al.
Publicado: (2025)
Ejemplares similares
-
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
por: Bai, Xu, et al.
Publicado: (2025) -
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
por: Bai, Haoyu, et al.
Publicado: (2026) -
DistributedEstimator: Distributed Training of Quantum Neural Networks via Circuit Cutting
por: Singh, Prabhjot, et al.
Publicado: (2026) -
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
por: Wen, Linfeng, et al.
Publicado: (2024) -
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
por: Bai, Haoyu, et al.
Publicado: (2024)