Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
Fuente:
arXiv
Saved in:
| Main Author: | Xie, Tianfang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
by: Chu, Zhengyan
Published: (2026)
by: Chu, Zhengyan
Published: (2026)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Shipwright: Proving liveness of distributed systems with Byzantine participants
by: Leung, Derek, et al.
Published: (2025)
by: Leung, Derek, et al.
Published: (2025)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)
by: Pan, Ding, et al.
Published: (2026)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
by: Khan, Shamsher
Published: (2026)
by: Khan, Shamsher
Published: (2026)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
HiDVFS: A Hierarchical Multi-Agent DVFS Scheduler for OpenMP DAG Workloads
by: Pivezhandi, Mohammad, et al.
Published: (2026)
by: Pivezhandi, Mohammad, et al.
Published: (2026)
Verifying In-Network Computing Systems for Design Risks
by: Bai, Tianyu, et al.
Published: (2026)
by: Bai, Tianyu, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads
by: Stokely, Murray, et al.
Published: (2025)
by: Stokely, Murray, et al.
Published: (2025)
Alea-BFT: Practical Asynchronous Byzantine Fault Tolerance
by: Antunes, Diogo S., et al.
Published: (2024)
by: Antunes, Diogo S., et al.
Published: (2024)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
by: Da, Wei, et al.
Published: (2025)
by: Da, Wei, et al.
Published: (2025)
Intelligent Cloud Orchestration: A Hybrid Predictive and Heuristic Framework for Cost Optimization
by: Nagoriya, Heet, et al.
Published: (2026)
by: Nagoriya, Heet, et al.
Published: (2026)
NimbusGuard: A Novel Framework for Proactive Kubernetes Autoscaling Using Deep Q-Networks
by: Wanigasooriya, Chamath, et al.
Published: (2026)
by: Wanigasooriya, Chamath, et al.
Published: (2026)
Semaphores Augmented with a Waiting Array
by: Dice, Dave, et al.
Published: (2025)
by: Dice, Dave, et al.
Published: (2025)
Reciprocating Locks
by: Dice, Dave, et al.
Published: (2025)
by: Dice, Dave, et al.
Published: (2025)
Hapax Locks : Value-Based Mutual Exclusion
by: Dice, Dave, et al.
Published: (2025)
by: Dice, Dave, et al.
Published: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
by: Hovland, Paul D.
Published: (2025)
by: Hovland, Paul D.
Published: (2025)
Using a Market Economy to Provision Compute Resources Across Planet-wide Clusters
by: Stokely, Murray, et al.
Published: (2025)
by: Stokely, Murray, et al.
Published: (2025)
Privacy-Aware Split Inference with Speculative Decoding for Large Language Models over Wide-Area Networks
by: Cunningham, Michael
Published: (2026)
by: Cunningham, Michael
Published: (2026)
Data Race Satisfiability on Array Elements
by: Shim, Junhyung, et al.
Published: (2025)
by: Shim, Junhyung, et al.
Published: (2025)
Big Data Workload Profiling for Energy-Aware Cloud Resource Management
by: Parikh, Milan, et al.
Published: (2026)
by: Parikh, Milan, et al.
Published: (2026)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
by: Park, Gyeongseo, et al.
Published: (2026)
by: Park, Gyeongseo, et al.
Published: (2026)
GPUnion: Autonomous GPU Sharing on Campus
by: Li, Yufang, et al.
Published: (2025)
by: Li, Yufang, et al.
Published: (2025)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
Distributed Recoverable Sketches (Extended Version)
by: Cohen, Diana, et al.
Published: (2025)
by: Cohen, Diana, et al.
Published: (2025)
Intersections of Web3 and AI -- View in 2024
by: Hyland-Wood, David, et al.
Published: (2024)
by: Hyland-Wood, David, et al.
Published: (2024)
Artifact Evaluation for Distributed Systems: Current Practices and Beyond
by: Sedghpour, Mohammad Reza Saleh, et al.
Published: (2024)
by: Sedghpour, Mohammad Reza Saleh, et al.
Published: (2024)
NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUs
by: Carver, Benjamin, et al.
Published: (2025)
by: Carver, Benjamin, et al.
Published: (2025)
Generic Multicast (Extended Version)
by: Bolina, José Augusto, et al.
Published: (2024)
by: Bolina, José Augusto, et al.
Published: (2024)
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
by: Ahmadvand, Mohsen, et al.
Published: (2026)
by: Ahmadvand, Mohsen, et al.
Published: (2026)
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
dpBento: Benchmarking DPUs for Data Processing
by: Hu, Jiasheng, et al.
Published: (2025)
by: Hu, Jiasheng, et al.
Published: (2025)
Parallel/Distributed Tabu Search for Scheduling Microprocessor Tasks in Hybrid Flowshop
by: Janiak, Adam, et al.
Published: (2025)
by: Janiak, Adam, et al.
Published: (2025)
Social Dynamics of DAOs: Power, Onboarding, and Inclusivity
by: Kozlova, Victoria, et al.
Published: (2025)
by: Kozlova, Victoria, et al.
Published: (2025)
Similar Items
-
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
by: Chu, Zhengyan
Published: (2026) -
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025) -
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025) -
Shipwright: Proving liveness of distributed systems with Byzantine participants
by: Leung, Derek, et al.
Published: (2025) -
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)