SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Dongxin, Wu, Jikun, Yiu, Siu Ming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)
by: Xu, Youxuan, et al.
Published: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023)
by: Chen, Jinfan, et al.
Published: (2023)
GRAIN: Exact Graph Reconstruction from Gradients
by: Drencheva, Maria, et al.
Published: (2025)
by: Drencheva, Maria, et al.
Published: (2025)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025)
by: Shi, Jinliang, et al.
Published: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
by: Li, Shigang, et al.
Published: (2022)
by: Li, Shigang, et al.
Published: (2022)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
by: Shi, Jinliang, et al.
Published: (2024)
by: Shi, Jinliang, et al.
Published: (2024)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
by: Li, Shigang, et al.
Published: (2021)
by: Li, Shigang, et al.
Published: (2021)
Securing Federated Sensitive Topic Classification against Poisoning Attacks
by: Chu, Tianyue, et al.
Published: (2022)
by: Chu, Tianyue, et al.
Published: (2022)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
by: Parakhin, Vladyslav
Published: (2026)
by: Parakhin, Vladyslav
Published: (2026)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Decentralized Task Scheduling in Distributed Systems: A Deep Reinforcement Learning Approach
by: John, Daniel Benniah
Published: (2026)
by: John, Daniel Benniah
Published: (2026)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Coalition Formation in LLM Agent Networks: Stability Analysis and Convergence Guarantees
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Learning to Collaborate: An Orchestrated-Decentralized Framework for Peer-to-Peer LLM Federation
by: Singh, Inderjeet, et al.
Published: (2026)
by: Singh, Inderjeet, et al.
Published: (2026)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
by: Hu, Zhibo, et al.
Published: (2024)
by: Hu, Zhibo, et al.
Published: (2024)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
by: Kang, Daemyung, et al.
Published: (2026)
by: Kang, Daemyung, et al.
Published: (2026)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
by: Zhao, Mingjie, et al.
Published: (2026)
by: Zhao, Mingjie, et al.
Published: (2026)
Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
by: Rao, Swanand, et al.
Published: (2026)
by: Rao, Swanand, et al.
Published: (2026)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
by: Petrella, Alfredo, et al.
Published: (2023)
by: Petrella, Alfredo, et al.
Published: (2023)
Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters
by: Sultana, Abeda, et al.
Published: (2025)
by: Sultana, Abeda, et al.
Published: (2025)
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
SDOF: Taming the Alignment Tax in Multi-Agent Orchestration with State-Constrained Dispatch
by: Wang, Zhantao
Published: (2026)
by: Wang, Zhantao
Published: (2026)
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
by: Chu, Zhengyan
Published: (2026)
by: Chu, Zhengyan
Published: (2026)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
by: Xie, Tianfang
Published: (2026)
by: Xie, Tianfang
Published: (2026)
Emergent Social Structures in Autonomous AI Agent Networks: A Metadata Analysis of 626 Agents on the Pilot Protocol
by: Calin, Teodor-Ioan
Published: (2026)
by: Calin, Teodor-Ioan
Published: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
by: Xiong, Yi, et al.
Published: (2024)
by: Xiong, Yi, et al.
Published: (2024)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
by: Parmar, Abhinav Singh
Published: (2026)
by: Parmar, Abhinav Singh
Published: (2026)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Semaphores Augmented with a Waiting Array
by: Dice, Dave, et al.
Published: (2025)
by: Dice, Dave, et al.
Published: (2025)
Reciprocating Locks
by: Dice, Dave, et al.
Published: (2025)
by: Dice, Dave, et al.
Published: (2025)
Similar Items
-
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025) -
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023) -
GRAIN: Exact Graph Reconstruction from Gradients
by: Drencheva, Maria, et al.
Published: (2025) -
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025) -
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)