Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
Fuente:
arXiv
Saved in:
| Main Authors: | McDaniel, Adam, Jantz, Michael, Sharma, Ashesh, Abbott, Steve, Martin, Steven, Khandekar, Shreyas, Neth, Brandon, Alvarez, Bruno Villasenor, Kashi, Aditya, Elwasif, Wael, Hernandez, Oscar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026)
by: Mhatre, Kaustubh, et al.
Published: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
by: Latif, Imran, et al.
Published: (2024)
by: Latif, Imran, et al.
Published: (2024)
AMD Versal Implementations of FAM and SSCA Estimators
by: Li, Carol Jingyi, et al.
Published: (2025)
by: Li, Carol Jingyi, et al.
Published: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
by: Singhania, Varsha, et al.
Published: (2024)
by: Singhania, Varsha, et al.
Published: (2024)
Scalable and Efficient Intra- and Inter-node Interconnection Networks for Post-Exascale Supercomputers and Data centers
by: Tarraga-Moreno, Joaquin, et al.
Published: (2025)
by: Tarraga-Moreno, Joaquin, et al.
Published: (2025)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
by: Rösti, André, et al.
Published: (2025)
by: Rösti, André, et al.
Published: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025)
by: Shin, Changmin, et al.
Published: (2025)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024)
by: Dai, Xilai, et al.
Published: (2024)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
VolTune: A Fine-Grained Runtime Voltage Control Architecture for FPGA Systems
by: Ahmed, Akram Ben, et al.
Published: (2026)
by: Ahmed, Akram Ben, et al.
Published: (2026)
RoboGPU: Accelerating GPU Collision Detection for Robotics
by: Liu, Lufei, et al.
Published: (2026)
by: Liu, Lufei, et al.
Published: (2026)
Annotating Slack Directly on Your Verilog: Fine-Grained RTL Timing Evaluation for Early Optimization
by: Fang, Wenji, et al.
Published: (2024)
by: Fang, Wenji, et al.
Published: (2024)
A Prototype-Based Framework to Design Scalable Heterogeneous SoCs with Fine-Grained DFS
by: Montanaro, Gabriele, et al.
Published: (2024)
by: Montanaro, Gabriele, et al.
Published: (2024)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
by: Olgun, Ataberk, et al.
Published: (2022)
by: Olgun, Ataberk, et al.
Published: (2022)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)
by: Langarita, Rubén, et al.
Published: (2025)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
by: Elwasif, Wael, et al.
Published: (2022)
by: Elwasif, Wael, et al.
Published: (2022)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
by: Makin, Yashasvi, et al.
Published: (2025)
by: Makin, Yashasvi, et al.
Published: (2025)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
by: Wang, Yimin, et al.
Published: (2025)
by: Wang, Yimin, et al.
Published: (2025)
FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian Splatting
by: Ou, Wenhui, et al.
Published: (2026)
by: Ou, Wenhui, et al.
Published: (2026)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
DeepAssert: An LLM-Aided Verification Framework with Fine-Grained Assertion Generation for Modules with Extracted Module Specifications
by: Wang, Yonghao, et al.
Published: (2025)
by: Wang, Yonghao, et al.
Published: (2025)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
Design of a GPU with Heterogeneous Cores for Graphics
by: Tomás, Aurora, et al.
Published: (2026)
by: Tomás, Aurora, et al.
Published: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024)
by: Luo, Weile, et al.
Published: (2024)
COOK Access Control on an embedded Volta GPU
by: Lesage, Benjamin, et al.
Published: (2024)
by: Lesage, Benjamin, et al.
Published: (2024)
Multiport Support for Vortex OpenGPU Memory Hierarchy
by: Shin, Injae, et al.
Published: (2025)
by: Shin, Injae, et al.
Published: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
by: Zhao, Jisheng, et al.
Published: (2026)
by: Zhao, Jisheng, et al.
Published: (2026)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
Latch Based Design for Fast Voltage Droop Response
by: Srinivas, Shreyas, et al.
Published: (2025)
by: Srinivas, Shreyas, et al.
Published: (2025)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
by: Cai, Victor, et al.
Published: (2025)
by: Cai, Victor, et al.
Published: (2025)
ATLAS: A Self-Supervised and Cross-Stage Netlist Power Model for Fine-Grained Time-Based Layout Power Analysis
by: Li, Wenkai, et al.
Published: (2025)
by: Li, Wenkai, et al.
Published: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
Automatic High-quality Verilog Assertion Generation through Subtask-Focused Fine-Tuned LLMs and Iterative Prompting
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
by: Nagendra, Savinay
Published: (2024)
by: Nagendra, Savinay
Published: (2024)
Similar Items
-
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026) -
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
by: Latif, Imran, et al.
Published: (2024) -
AMD Versal Implementations of FAM and SSCA Estimators
by: Li, Carol Jingyi, et al.
Published: (2025) -
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
by: Singhania, Varsha, et al.
Published: (2024) -
Scalable and Efficient Intra- and Inter-node Interconnection Networks for Post-Exascale Supercomputers and Data centers
by: Tarraga-Moreno, Joaquin, et al.
Published: (2025)