Online GPU Energy Optimization with Switching-Aware Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xiongxiao, Bekele, Solomon Abera, Videau, Brice, Shu, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
by: Park, Jeongmin Brian, et al.
Published: (2023)
by: Park, Jeongmin Brian, et al.
Published: (2023)
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
by: Pang, Bowen, et al.
Published: (2025)
by: Pang, Bowen, et al.
Published: (2025)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
by: Makin, Yashasvi, et al.
Published: (2025)
by: Makin, Yashasvi, et al.
Published: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System
by: Rhyner, Steve, et al.
Published: (2024)
by: Rhyner, Steve, et al.
Published: (2024)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
by: Zhao, Dan, et al.
Published: (2024)
by: Zhao, Dan, et al.
Published: (2024)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
by: Dorairaj, Nij, et al.
Published: (2026)
by: Dorairaj, Nij, et al.
Published: (2026)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
by: Panigrahy, Deepak, et al.
Published: (2026)
by: Panigrahy, Deepak, et al.
Published: (2026)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
by: Zhu, Wenbin, et al.
Published: (2025)
by: Zhu, Wenbin, et al.
Published: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
by: Chen, Jiesong, et al.
Published: (2026)
by: Chen, Jiesong, et al.
Published: (2026)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025)
by: Choudhary, Mansi, et al.
Published: (2025)
Serving Large Language Models on Huawei CloudMatrix384
by: Zuo, Pengfei, et al.
Published: (2025)
by: Zuo, Pengfei, et al.
Published: (2025)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
by: Shi, Tianyao, et al.
Published: (2025)
by: Shi, Tianyao, et al.
Published: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
ZettaLith: An Architectural Exploration of Extreme-Scale AI Inference Acceleration
by: Silverbrook, Kia
Published: (2025)
by: Silverbrook, Kia
Published: (2025)
Enabling Accelerators for Graph Computing
by: Shivdikar, Kaustubh
Published: (2023)
by: Shivdikar, Kaustubh
Published: (2023)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
by: Yu, Zhongkai, et al.
Published: (2025)
by: Yu, Zhongkai, et al.
Published: (2025)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
by: Yu, Zhongkai, et al.
Published: (2026)
by: Yu, Zhongkai, et al.
Published: (2026)
Splitwiser: Efficient LM inference with constrained resources
by: Aali, Asad, et al.
Published: (2025)
by: Aali, Asad, et al.
Published: (2025)
CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead
by: Ye, Jinpeng, et al.
Published: (2026)
by: Ye, Jinpeng, et al.
Published: (2026)
Sensitivity-Guided Framework for Pruned and Quantized Reservoir Computing Accelerators
by: Jafari, Atousa, et al.
Published: (2026)
by: Jafari, Atousa, et al.
Published: (2026)
PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System
by: He, Yintao, et al.
Published: (2025)
by: He, Yintao, et al.
Published: (2025)
FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
by: Shen, Ao, et al.
Published: (2025)
by: Shen, Ao, et al.
Published: (2025)
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
by: Choi, Yuseon, et al.
Published: (2026)
by: Choi, Yuseon, et al.
Published: (2026)
VLSI Hypergraph Partitioning with Deep Learning
by: Khan, Muhammad Hadir, et al.
Published: (2024)
by: Khan, Muhammad Hadir, et al.
Published: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
A High Energy-Efficiency Multi-core Neuromorphic Architecture for Deep SNN Training
by: Li, Mingjing, et al.
Published: (2024)
by: Li, Mingjing, et al.
Published: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
by: DeBole, Michael V., et al.
Published: (2025)
by: DeBole, Michael V., et al.
Published: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
by: Blanco, Francesco G., et al.
Published: (2024)
by: Blanco, Francesco G., et al.
Published: (2024)
Similar Items
-
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024) -
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
by: Park, Jeongmin Brian, et al.
Published: (2023) -
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
by: Pang, Bowen, et al.
Published: (2025) -
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
by: Makin, Yashasvi, et al.
Published: (2025) -
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)