FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zirui, Hu, Yi-Xiang, Wu, Feng, Li, Xiangyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
by: Hu, Yi-Xiang, et al.
Published: (2026)
by: Hu, Yi-Xiang, et al.
Published: (2026)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
by: Kulagina, Svetlana, et al.
Published: (2025)
by: Kulagina, Svetlana, et al.
Published: (2025)
WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2025)
by: Lehmann, Fabian, et al.
Published: (2025)
Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows
by: Schweisgut, Dominik, et al.
Published: (2026)
by: Schweisgut, Dominik, et al.
Published: (2026)
Carbon-Aware Workflow Scheduling with Fixed Mapping and Deadline Constraint
by: Schweisgut, Dominik, et al.
Published: (2025)
by: Schweisgut, Dominik, et al.
Published: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)
by: Pagonas, Nikos, et al.
Published: (2025)
Past-Future Scheduler for LLM Serving under SLA Guarantees
by: Gong, Ruihao, et al.
Published: (2025)
by: Gong, Ruihao, et al.
Published: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
by: Russo, Gabriele Russo, et al.
Published: (2024)
by: Russo, Gabriele Russo, et al.
Published: (2024)
Reinforcement Learning based Workflow Scheduling in Cloud and Edge Computing Environments: A Taxonomy, Review and Future Directions
by: Jayanetti, Amanda, et al.
Published: (2024)
by: Jayanetti, Amanda, et al.
Published: (2024)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
by: Idouar, Yacine, et al.
Published: (2025)
by: Idouar, Yacine, et al.
Published: (2025)
Workflows Community Summit 2024: Future Trends and Challenges in Scientific Workflows
by: da Silva, Rafael Ferreira, et al.
Published: (2024)
by: da Silva, Rafael Ferreira, et al.
Published: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration
by: Guan, Wen, et al.
Published: (2025)
by: Guan, Wen, et al.
Published: (2025)
Minimizing Energy in Reliability and Deadline-Ensured Workflow Scheduling in Cloud
by: Sarkar, Suvarthi, et al.
Published: (2025)
by: Sarkar, Suvarthi, et al.
Published: (2025)
sAirflow: Adopting Serverless in a Legacy Workflow Scheduler
by: Mikina, Filip, et al.
Published: (2024)
by: Mikina, Filip, et al.
Published: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
by: Loi, Chung Ming, et al.
Published: (2025)
by: Loi, Chung Ming, et al.
Published: (2025)
SLURM Heterogeneous Jobs for Hybrid Classical-Quantum Workflows
by: Esposito, Aniello, et al.
Published: (2025)
by: Esposito, Aniello, et al.
Published: (2025)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
by: Wei, Xinming, et al.
Published: (2025)
by: Wei, Xinming, et al.
Published: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
Scientific Workflow Scheduling in Cloud Considering Cold Start and Variable Pricing Model
by: Sarkar, Suvarthi, et al.
Published: (2025)
by: Sarkar, Suvarthi, et al.
Published: (2025)
Enabling Scientific Workflow Scheduling Research in Non-Uniform Memory Access Architectures
by: Vivas, Aurelio, et al.
Published: (2025)
by: Vivas, Aurelio, et al.
Published: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Mapping Large Memory-constrained Workflows onto Heterogeneous Platforms
by: Kulagina, Svetlana, et al.
Published: (2024)
by: Kulagina, Svetlana, et al.
Published: (2024)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
by: Li, Pengbo, et al.
Published: (2026)
by: Li, Pengbo, et al.
Published: (2026)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
by: Chang, Zihan, et al.
Published: (2024)
by: Chang, Zihan, et al.
Published: (2024)
Workflow-Driven Modeling for the Compute Continuum: An Optimization Approach to Automated System and Workload Scheduling
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
by: Shan, Chenggang, et al.
Published: (2024)
by: Shan, Chenggang, et al.
Published: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
by: Hui, Xinning, et al.
Published: (2024)
by: Hui, Xinning, et al.
Published: (2024)
Augur: Pre-Execution Energy Prediction for Workflow Tasks in Heterogeneous Clusters
by: West, Kathleen, et al.
Published: (2026)
by: West, Kathleen, et al.
Published: (2026)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
Exploring the Potential of Carbon-Aware Execution for Scientific Workflows
by: West, Kathleen, et al.
Published: (2025)
by: West, Kathleen, et al.
Published: (2025)
Similar Items
-
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026) -
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
by: Hu, Yi-Xiang, et al.
Published: (2026) -
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
by: Kulagina, Svetlana, et al.
Published: (2025) -
WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2025) -
Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows
by: Schweisgut, Dominik, et al.
Published: (2026)