Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Weile, Fan, Ruibo, Li, Zeyu, Du, Dayou, Liu, Hongyuan, Wang, Qiang, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024)
von: Luo, Weile, et al.
Veröffentlicht: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
von: Jiang, Wenqi
Veröffentlicht: (2025)
von: Jiang, Wenqi
Veröffentlicht: (2025)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
von: Tang, Xinsheng, et al.
Veröffentlicht: (2026)
von: Tang, Xinsheng, et al.
Veröffentlicht: (2026)
On Optimizing Locality of Graph Transposition on Modern Architectures
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2025)
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2025)
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
Ultra Ethernet's Design Principles and Architectural Innovations
von: Hoefler, Torsten, et al.
Veröffentlicht: (2025)
von: Hoefler, Torsten, et al.
Veröffentlicht: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
Random Adaptive Cache Placement Policy
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026) -
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025) -
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024) -
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025) -
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)