An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ting, Hsu-Tzu, Chou, Jerry, Chen, Ming-Hung, Chung, I-Hsin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
von: Zambianco, Marco, et al.
Veröffentlicht: (2025)
von: Zambianco, Marco, et al.
Veröffentlicht: (2025)
Managing Multi Instance GPUs for High Throughput and Energy Savings
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
Night-Window Batching versus Carbon-Aware Scheduling for Clinical AI GPU Workloads
von: Doshi, Nishi, et al.
Veröffentlicht: (2026)
von: Doshi, Nishi, et al.
Veröffentlicht: (2026)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
Learning to Schedule: A Supervised Learning Framework for Network-Aware Scheduling of Data-Intensive Workloads
von: Timilsina, Sankalpa, et al.
Veröffentlicht: (2025)
von: Timilsina, Sankalpa, et al.
Veröffentlicht: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
Engineering A Workload-balanced Push-Relabel Algorithm for Massive Graphs on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
SWARM+: Scalable and Resilient Multi-Agent Consensus for Fully-Decentralized Data-Aware Workload Management
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
On the Partitioning of GPU Power among Multi-Instances
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
Guardian: Safe GPU Sharing in Multi-Tenant Environments
von: Pavlidakis, Manos, et al.
Veröffentlicht: (2024)
von: Pavlidakis, Manos, et al.
Veröffentlicht: (2024)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
Eventually-Consistent Federated Scheduling for Data Center Workloads
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2023)
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2023)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
GPU Sharing with Triples Mode
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
Cuckoo-GPU: Accelerating Cuckoo Filters on Modern GPUs
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
von: Wu, Xin, et al.
Veröffentlicht: (2026)
von: Wu, Xin, et al.
Veröffentlicht: (2026)
Data Management System Analysis for Distributed Computing Workloads
von: Hsu, Kuan-Chieh, et al.
Veröffentlicht: (2025)
von: Hsu, Kuan-Chieh, et al.
Veröffentlicht: (2025)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2025)
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2025)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
von: Scheinert, Dominik, et al.
Veröffentlicht: (2026)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2026)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024) -
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2024) -
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
von: Zambianco, Marco, et al.
Veröffentlicht: (2025) -
Managing Multi Instance GPUs for High Throughput and Energy Savings
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025) -
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)