Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Lingling, Zhang, Gen, Peng, Jialin, Xu, Xiang, Xu, Yuan, Ma, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025)
by: Khan, Kabir, et al.
Published: (2025)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
by: Sunkara, Krishna Chaitanya
Published: (2026)
by: Sunkara, Krishna Chaitanya
Published: (2026)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
by: Yang, Haotian, et al.
Published: (2025)
by: Yang, Haotian, et al.
Published: (2025)
ACME: Adaptive Customization of Large Models via Distributed Systems
by: Dai, Ziming, et al.
Published: (2025)
by: Dai, Ziming, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Knowledge Graphs-Driven Intelligence for Distributed Decision Systems
by: Napoli, Rosario, et al.
Published: (2026)
by: Napoli, Rosario, et al.
Published: (2026)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
by: Rosendal, Daan, et al.
Published: (2026)
by: Rosendal, Daan, et al.
Published: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
by: Zhang, Guilin, et al.
Published: (2026)
by: Zhang, Guilin, et al.
Published: (2026)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Federated Few-Shot Learning on Neuromorphic Hardware: An Empirical Study Across Physical Edge Nodes
by: Motta, Steven, et al.
Published: (2026)
by: Motta, Steven, et al.
Published: (2026)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
A Selective Homomorphic Encryption Approach for Faster Privacy-Preserving Federated Learning
by: Korkmaz, Abdulkadir, et al.
Published: (2025)
by: Korkmaz, Abdulkadir, et al.
Published: (2025)
SepsisAI Orchestrator: A Containerized and Scalable Platform for Deploying AI Models and Real-Time Monitoring in Early Sepsis Detection
by: Ospitia, Santiago, et al.
Published: (2026)
by: Ospitia, Santiago, et al.
Published: (2026)
SRFed: Mitigating Poisoning Attacks in Privacy-Preserving Federated Learning with Heterogeneous Data
by: Lu, Yiwen
Published: (2026)
by: Lu, Yiwen
Published: (2026)
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
by: Saleh, Alaa, et al.
Published: (2023)
by: Saleh, Alaa, et al.
Published: (2023)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
by: Da, Wei, et al.
Published: (2025)
by: Da, Wei, et al.
Published: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
by: Petrella, Alfredo, et al.
Published: (2023)
by: Petrella, Alfredo, et al.
Published: (2023)
Intelligent Cloud Orchestration: A Hybrid Predictive and Heuristic Framework for Cost Optimization
by: Nagoriya, Heet, et al.
Published: (2026)
by: Nagoriya, Heet, et al.
Published: (2026)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
De-DSI: Decentralised Differentiable Search Index
by: Neague, Petru, et al.
Published: (2024)
by: Neague, Petru, et al.
Published: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026)
by: Topcu, Burak, et al.
Published: (2026)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
by: Nouri, Azam
Published: (2026)
by: Nouri, Azam
Published: (2026)
MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
by: Mao, Wei-Lung, et al.
Published: (2025)
by: Mao, Wei-Lung, et al.
Published: (2025)
Network Structures as an Attack Surface: Topology-Based Privacy Leakage in Federated Learning
by: Rangwala, Murtaza, et al.
Published: (2025)
by: Rangwala, Murtaza, et al.
Published: (2025)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
by: Karabulut, Burak, et al.
Published: (2026)
by: Karabulut, Burak, et al.
Published: (2026)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
Similar Items
-
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025) -
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
by: Sunkara, Krishna Chaitanya
Published: (2026) -
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025) -
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025) -
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)