Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Jarmusch, Aaron, Graddon, Nathan, Chandrasekaran, Sunita |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
by: Jarmusch, Aaron, et al.
Published: (2025)
by: Jarmusch, Aaron, et al.
Published: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
by: Liu, Rixin, et al.
Published: (2026)
by: Liu, Rixin, et al.
Published: (2026)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024)
by: Tramm, John, et al.
Published: (2024)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
by: Vanecek, Stepan, et al.
Published: (2025)
by: Vanecek, Stepan, et al.
Published: (2025)
Supercharging Federated Learning with Flower and NVIDIA FLARE
by: Roth, Holger R., et al.
Published: (2024)
by: Roth, Holger R., et al.
Published: (2024)
Empowering Federated Learning for Massive Models with NVIDIA FLARE
by: Roth, Holger R., et al.
Published: (2024)
by: Roth, Holger R., et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
by: Bikshandi, Ganesh
Published: (2026)
by: Bikshandi, Ganesh
Published: (2026)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
by: Raffin, Guillaume, et al.
Published: (2024)
by: Raffin, Guillaume, et al.
Published: (2024)
Big Data Architecture for Large Organizations
by: Ismail, Fathima Nuzla, et al.
Published: (2025)
by: Ismail, Fathima Nuzla, et al.
Published: (2025)
Serverless Computing: Architecture, Concepts, and Applications
by: Ghorbian, Mohsen, et al.
Published: (2025)
by: Ghorbian, Mohsen, et al.
Published: (2025)
Resolving Conflicts with Grace: Dynamically Concurrent Universality
by: Kuznetsov, Petr, et al.
Published: (2025)
by: Kuznetsov, Petr, et al.
Published: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
Optimizing Bloom Filters for Modern GPU Architectures
by: Jünger, Daniel, et al.
Published: (2025)
by: Jünger, Daniel, et al.
Published: (2025)
Can you keep a secret? A new protocol for sender-side enforcement of causal message delivery
by: Tong, Yan, et al.
Published: (2026)
by: Tong, Yan, et al.
Published: (2026)
Dependency-Aware Execution Mechanism in Hyperledger Fabric Architecture
by: Kaul, Sanyam, et al.
Published: (2025)
by: Kaul, Sanyam, et al.
Published: (2025)
FOGNITE: Federated Learning-Enhanced Fog-Cloud Architecture
by: Sobati-M, Somayeh
Published: (2025)
by: Sobati-M, Somayeh
Published: (2025)
Mapping Gemma3 onto an Edge Dataflow Architecture
by: Du, Shouyu, et al.
Published: (2026)
by: Du, Shouyu, et al.
Published: (2026)
Comparative Study of Large Language Model Architectures on Frontier
by: Yin, Junqi, et al.
Published: (2024)
by: Yin, Junqi, et al.
Published: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
by: Ahmad, Hussain, et al.
Published: (2025)
by: Ahmad, Hussain, et al.
Published: (2025)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
by: Kulagina, Svetlana, et al.
Published: (2025)
by: Kulagina, Svetlana, et al.
Published: (2025)
Assessing Redundancy Strategies to Improve Availability in Virtualized System Architectures
by: Silva, Alison, et al.
Published: (2025)
by: Silva, Alison, et al.
Published: (2025)
Modular Architecture for High-Performance and Low Overhead Data Transfers
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
Data Augmentation and Convolutional Network Architecture Influence on Distributed Learning
by: Jansen, Victor Forattini, et al.
Published: (2026)
by: Jansen, Victor Forattini, et al.
Published: (2026)
Akita: A High Usability Simulation Framework for Computer Architecture
by: Jannat, Sabila Al, et al.
Published: (2026)
by: Jannat, Sabila Al, et al.
Published: (2026)
Spider: A BFT Architecture for Geo-Replicated Cloud Services
by: Eischer, Michael, et al.
Published: (2024)
by: Eischer, Michael, et al.
Published: (2024)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
by: Giannoula, Christina, et al.
Published: (2024)
by: Giannoula, Christina, et al.
Published: (2024)
Hybrid Cloud Architectures for Research Computing: Applications and Use Cases
by: Stiensmeier, Xaver, et al.
Published: (2026)
by: Stiensmeier, Xaver, et al.
Published: (2026)
Benchmarking Machine Learning Applications on Heterogeneous Architecture using Reframe
by: Rae, Christopher, et al.
Published: (2024)
by: Rae, Christopher, et al.
Published: (2024)
Similar Items
-
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026) -
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
by: Jarmusch, Aaron, et al.
Published: (2026) -
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025) -
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
by: Jarmusch, Aaron, et al.
Published: (2025) -
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
by: Li, Junjie, et al.
Published: (2024)