Optimal Workload Placement on Multi-Instance GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Turkkan, Bekir, Murali, Pavankumar, Harsha, Pavithra, Arora, Rohan, Vanloo, Gerard, Narayanaswami, Chandra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
Managing Multi Instance GPUs for High Throughput and Energy Savings
di: Saraha, Abhijeet, et al.
Pubblicazione: (2025)
di: Saraha, Abhijeet, et al.
Pubblicazione: (2025)
How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
di: Cui, Shengkun, et al.
Pubblicazione: (2025)
di: Cui, Shengkun, et al.
Pubblicazione: (2025)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
di: Siavashi, Ahmad, et al.
Pubblicazione: (2025)
di: Siavashi, Ahmad, et al.
Pubblicazione: (2025)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
di: Li, Zhilin, et al.
Pubblicazione: (2025)
di: Li, Zhilin, et al.
Pubblicazione: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
Accelerating Maximal Biclique Enumeration on GPUs
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
Optimizing sDTW for AMD GPUs
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
Agentic AI Workload Characteristics
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
Engineering A Workload-balanced Push-Relabel Algorithm for Massive Graphs on GPUs
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
di: Liu, Jinyang, et al.
Pubblicazione: (2023)
di: Liu, Jinyang, et al.
Pubblicazione: (2023)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
SWARM+: Scalable and Resilient Multi-Agent Consensus for Fully-Decentralized Data-Aware Workload Management
di: Thareja, Komal, et al.
Pubblicazione: (2026)
di: Thareja, Komal, et al.
Pubblicazione: (2026)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
di: Wang, Tianyu, et al.
Pubblicazione: (2024)
di: Wang, Tianyu, et al.
Pubblicazione: (2024)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
di: Plesner, Andreas, et al.
Pubblicazione: (2024)
di: Plesner, Andreas, et al.
Pubblicazione: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
di: Chen, Yinfang, et al.
Pubblicazione: (2025)
di: Chen, Yinfang, et al.
Pubblicazione: (2025)
AI Surrogate Model for Distributed Computing Workloads
di: Park, David K., et al.
Pubblicazione: (2024)
di: Park, David K., et al.
Pubblicazione: (2024)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
di: Stoyanov, Radostin, et al.
Pubblicazione: (2025)
di: Stoyanov, Radostin, et al.
Pubblicazione: (2025)
Accelerating Compound LLM Training Workloads with Maestro
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
Union: An Automatic Workload Manager for Accelerating Network Simulation
di: Wang, Xin, et al.
Pubblicazione: (2024)
di: Wang, Xin, et al.
Pubblicazione: (2024)
Distributed Load Balancing with Workload-Dependent Service Rates
di: Zhang, Wenxin, et al.
Pubblicazione: (2024)
di: Zhang, Wenxin, et al.
Pubblicazione: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
di: Medeiros, Daniel, et al.
Pubblicazione: (2024)
di: Medeiros, Daniel, et al.
Pubblicazione: (2024)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
di: Huang, Lexiang, et al.
Pubblicazione: (2024)
di: Huang, Lexiang, et al.
Pubblicazione: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025) -
Managing Multi Instance GPUs for High Throughput and Energy Savings
di: Saraha, Abhijeet, et al.
Pubblicazione: (2025) -
How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis
di: Turkkan, Bekir, et al.
Pubblicazione: (2024) -
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
di: Cui, Shengkun, et al.
Pubblicazione: (2025) -
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
di: Siavashi, Ahmad, et al.
Pubblicazione: (2025)