Managing Multi Instance GPUs for High Throughput and Energy Savings
Fuente:
arXiv
Saved in:
| Main Authors: | Saraha, Abhijeet, Li, Yuanbo, Porter, Chris, Pande, Santosh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
by: Ting, Hsu-Tzu, et al.
Published: (2025)
by: Ting, Hsu-Tzu, et al.
Published: (2025)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
by: Liu, Jinyang, et al.
Published: (2023)
by: Liu, Jinyang, et al.
Published: (2023)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
by: Pande, Sohan Kumar, et al.
Published: (2025)
by: Pande, Sohan Kumar, et al.
Published: (2025)
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
by: Karim, Abrarul, et al.
Published: (2026)
by: Karim, Abrarul, et al.
Published: (2026)
Beyond BFS: A Comparative Study of Rooted Spanning Tree Algorithms on GPUs
by: Sahu, Abhijeet, et al.
Published: (2026)
by: Sahu, Abhijeet, et al.
Published: (2026)
Mazzaroth: A High-Throughput DAG Consensus with State Root
by: Li, Haohan
Published: (2025)
by: Li, Haohan
Published: (2025)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
by: León-Vega, Luis G., et al.
Published: (2024)
by: León-Vega, Luis G., et al.
Published: (2024)
Shoal++: High Throughput DAG BFT Can Be Fast!
by: Arun, Balaji, et al.
Published: (2024)
by: Arun, Balaji, et al.
Published: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs
by: Li, Yanliang, et al.
Published: (2025)
by: Li, Yanliang, et al.
Published: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Analysis of Server Throughput For Managed Big Data Analytics Frameworks
by: Anagnostakis, Emmanouil, et al.
Published: (2025)
by: Anagnostakis, Emmanouil, et al.
Published: (2025)
Energy-Aware Workflow Execution: An Overview of Techniques for Saving Energy and Emissions in Scientific Compute Clusters
by: Thamsen, Lauritz, et al.
Published: (2025)
by: Thamsen, Lauritz, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
by: Pan, Qilong, et al.
Published: (2025)
by: Pan, Qilong, et al.
Published: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Optimizing STAR Aligner for High Throughput Computing in the Cloud
by: Kica, Piotr, et al.
Published: (2024)
by: Kica, Piotr, et al.
Published: (2024)
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
by: Ye, Haojie, et al.
Published: (2026)
by: Ye, Haojie, et al.
Published: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
by: Kim, Sukjin, et al.
Published: (2025)
by: Kim, Sukjin, et al.
Published: (2025)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
by: Chen, Xing, et al.
Published: (2025)
by: Chen, Xing, et al.
Published: (2025)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
by: Hidayetoglu, Mert, et al.
Published: (2025)
by: Hidayetoglu, Mert, et al.
Published: (2025)
BBCA-CHAIN: Low Latency, High Throughput BFT Consensus on a DAG
by: Malkhi, Dahlia, et al.
Published: (2023)
by: Malkhi, Dahlia, et al.
Published: (2023)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
by: Mittal, Kashish, et al.
Published: (2026)
by: Mittal, Kashish, et al.
Published: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
by: Plesner, Andreas, et al.
Published: (2024)
by: Plesner, Andreas, et al.
Published: (2024)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Similar Items
-
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024) -
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
by: Ting, Hsu-Tzu, et al.
Published: (2025) -
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
by: Liu, Jinyang, et al.
Published: (2023) -
An Efficient Approach for Energy Conservation in Cloud Computing Environment
by: Pande, Sohan Kumar, et al.
Published: (2025) -
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
by: Karim, Abrarul, et al.
Published: (2026)