Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Weile, Fan, Ruibo, Li, Zeyu, Du, Dayou, Wang, Qiang, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
von: Kundu, Yudhishthira, et al.
Veröffentlicht: (2025)
von: Kundu, Yudhishthira, et al.
Veröffentlicht: (2025)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024)
von: Du, Dayou, et al.
Veröffentlicht: (2024)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
von: Nagendra, Savinay
Veröffentlicht: (2024)
von: Nagendra, Savinay
Veröffentlicht: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
Dissecting Conditional Branch Predictors of Apple Firestorm and Qualcomm Oryon for Software Optimization and Architectural Analysis
von: Chen, Jiajie, et al.
Veröffentlicht: (2024)
von: Chen, Jiajie, et al.
Veröffentlicht: (2024)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
von: Pei, Minnan, et al.
Veröffentlicht: (2025)
von: Pei, Minnan, et al.
Veröffentlicht: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
von: Guo, Zeyu
Veröffentlicht: (2025)
von: Guo, Zeyu
Veröffentlicht: (2025)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
Survey on Characterizing and Understanding GNNs from a Computer Architecture Perspective
von: Wu, Meng, et al.
Veröffentlicht: (2024)
von: Wu, Meng, et al.
Veröffentlicht: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
RoboGPU: Accelerating GPU Collision Detection for Robotics
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
COOK Access Control on an embedded Volta GPU
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
Design of a GPU with Heterogeneous Cores for Graphics
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures
von: Li, Shangkun, et al.
Veröffentlicht: (2026)
von: Li, Shangkun, et al.
Veröffentlicht: (2026)
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
von: Zhou, Yaoyun, et al.
Veröffentlicht: (2025)
von: Zhou, Yaoyun, et al.
Veröffentlicht: (2025)
Adaptive Hybrid FFT: A Novel Pipeline and Memory-Based Architecture for Radix-$2^k$ FFT in Large Size Processing
von: Zhao, Fangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Fangyu, et al.
Veröffentlicht: (2025)
Multiport Support for Vortex OpenGPU Memory Hierarchy
von: Shin, Injae, et al.
Veröffentlicht: (2025)
von: Shin, Injae, et al.
Veröffentlicht: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
von: Dai, Tuo, et al.
Veröffentlicht: (2024)
von: Dai, Tuo, et al.
Veröffentlicht: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
von: Guan, Yue, et al.
Veröffentlicht: (2026)
von: Guan, Yue, et al.
Veröffentlicht: (2026)
Overmind NSA: A Unified Neuro-Symbolic Computing Architecture with Approximate Nonlinear Activations and Preemptive Memory Bypass
von: Wang, Weilun, et al.
Veröffentlicht: (2026)
von: Wang, Weilun, et al.
Veröffentlicht: (2026)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
Reconfigurable Stream Network Architecture
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
von: Latif, Imran, et al.
Veröffentlicht: (2024)
von: Latif, Imran, et al.
Veröffentlicht: (2024)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025) -
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026) -
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
von: Kundu, Yudhishthira, et al.
Veröffentlicht: (2025) -
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024) -
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025)