Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Saroliya, Urvij, Arima, Eishi, Liu, Dai, Schulz, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025)
by: Choudhary, Mansi, et al.
Published: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
by: Blanco, Francesco G., et al.
Published: (2024)
by: Blanco, Francesco G., et al.
Published: (2024)
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
by: Russo, Enrico, et al.
Published: (2024)
by: Russo, Enrico, et al.
Published: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
VLSI Hypergraph Partitioning with Deep Learning
by: Khan, Muhammad Hadir, et al.
Published: (2024)
by: Khan, Muhammad Hadir, et al.
Published: (2024)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
by: Yeo, Gwangoo, et al.
Published: (2026)
by: Yeo, Gwangoo, et al.
Published: (2026)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
by: Shi, Tianyao, et al.
Published: (2024)
by: Shi, Tianyao, et al.
Published: (2024)
HetGPU: The pursuit of making binary compatibility towards GPUs
by: Yang, Yiwei, et al.
Published: (2025)
by: Yang, Yiwei, et al.
Published: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
by: Liu, Yiqi, et al.
Published: (2025)
by: Liu, Yiqi, et al.
Published: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
Exploration of Cryptocurrency Mining-Specific GPUs in AI Applications: A Case Study of CMP 170HX
by: Kangwei, Xing
Published: (2025)
by: Kangwei, Xing
Published: (2025)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
by: Lee, You Hak, et al.
Published: (2025)
by: Lee, You Hak, et al.
Published: (2025)
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
by: Centeno, Sed, et al.
Published: (2025)
by: Centeno, Sed, et al.
Published: (2025)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
by: Yang, Jinwu, et al.
Published: (2026)
by: Yang, Jinwu, et al.
Published: (2026)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
by: Tschand, Arya, et al.
Published: (2024)
by: Tschand, Arya, et al.
Published: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023)
by: Hsia, Samuel, et al.
Published: (2023)
A Modern Primer on Processing in Memory
by: Mutlu, Onur, et al.
Published: (2020)
by: Mutlu, Onur, et al.
Published: (2020)
Memory-Aware Partitioning of Machine Learning Applications for Optimal Energy Use in Batteryless Systems
by: Gomez, Andres, et al.
Published: (2021)
by: Gomez, Andres, et al.
Published: (2021)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
by: Saba, Issa, et al.
Published: (2024)
by: Saba, Issa, et al.
Published: (2024)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
by: Jobst, Matthias, et al.
Published: (2025)
by: Jobst, Matthias, et al.
Published: (2025)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
by: Kreß, Fabian, et al.
Published: (2024)
by: Kreß, Fabian, et al.
Published: (2024)
A Survey on Graph Neural Network Acceleration: Algorithms, Systems, and Customized Hardware
by: Zhang, Shichang, et al.
Published: (2023)
by: Zhang, Shichang, et al.
Published: (2023)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
A High Energy-Efficiency Multi-core Neuromorphic Architecture for Deep SNN Training
by: Li, Mingjing, et al.
Published: (2024)
by: Li, Mingjing, et al.
Published: (2024)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
by: Abi-Karam, Stefan, et al.
Published: (2023)
by: Abi-Karam, Stefan, et al.
Published: (2023)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
Efficient, VRAM-Constrained xLM Inference on Clients
by: Ukarande, Aditya, et al.
Published: (2026)
by: Ukarande, Aditya, et al.
Published: (2026)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
Llumnix: Dynamic Scheduling for Large Language Model Serving
by: Sun, Biao, et al.
Published: (2024)
by: Sun, Biao, et al.
Published: (2024)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
by: Ding, Jianru, et al.
Published: (2026)
by: Ding, Jianru, et al.
Published: (2026)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
by: Zhang, Hengrui, et al.
Published: (2025)
by: Zhang, Hengrui, et al.
Published: (2025)
WWW: What, When, Where to Compute-in-Memory
by: Sharma, Tanvi, et al.
Published: (2023)
by: Sharma, Tanvi, et al.
Published: (2023)
AI and Memory Wall
by: Gholami, Amir, et al.
Published: (2024)
by: Gholami, Amir, et al.
Published: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
by: Russo, Enrico, et al.
Published: (2026)
by: Russo, Enrico, et al.
Published: (2026)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
Similar Items
-
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025) -
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025) -
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
by: Blanco, Francesco G., et al.
Published: (2024) -
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
by: Russo, Enrico, et al.
Published: (2024) -
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)