Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Shengkun, Patke, Archit, Nguyen, Hung, Ranjan, Aditya, Chen, Ziheng, Cao, Phuong, Bauer, Gregory, Bode, Brett, Di Martino, Catello, Jha, Saurabh, Narayanaswami, Chandra, Sow, Daby, Kalbarczyk, Zbigniew T., Iyer, Ravishankar K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
Queue management for slo-oriented large language model serving
by: Patke, Archit, et al.
Published: (2024)
by: Patke, Archit, et al.
Published: (2024)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
by: Qiu, Haoran, et al.
Published: (2024)
by: Qiu, Haoran, et al.
Published: (2024)
Security Testbed for Preempting Attacks against Supercomputing Infrastructure
by: Cao, Phuong, et al.
Published: (2024)
by: Cao, Phuong, et al.
Published: (2024)
Mutiny! How does Kubernetes fail, and what can we do about it?
by: Barletta, Marco, et al.
Published: (2024)
by: Barletta, Marco, et al.
Published: (2024)
PRAXIS: Integrating Program Analysis with Observability for Root-Cause Analysis
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
Collective Communication for 100k+ GPUs
by: Si, Min, et al.
Published: (2025)
by: Si, Min, et al.
Published: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
by: Yadav, Divakar Kumar, et al.
Published: (2026)
by: Yadav, Divakar Kumar, et al.
Published: (2026)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
by: Salpekar, Omkar, et al.
Published: (2026)
by: Salpekar, Omkar, et al.
Published: (2026)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
by: Zhang, Wuyue, et al.
Published: (2026)
by: Zhang, Wuyue, et al.
Published: (2026)
Efficient Coupled-Cluster Python Frameworks for Next-Generation GPUs: A Comparative Study of CuPy and PyTorch on the Hopper and Grace Hopper Architecture
by: Dobrowolska, Antonina, et al.
Published: (2026)
by: Dobrowolska, Antonina, et al.
Published: (2026)
Task-Based Tensor Computations on Modern GPUs
by: Yadav, Rohan, et al.
Published: (2025)
by: Yadav, Rohan, et al.
Published: (2025)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
by: Lee, Jin, et al.
Published: (2026)
by: Lee, Jin, et al.
Published: (2026)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
LTL learning on GPUs
by: Valizadeh, Mojtaba, et al.
Published: (2024)
by: Valizadeh, Mojtaba, et al.
Published: (2024)
A Context Engineering Framework for Improving Enterprise AI Agents based on Digital-Twin MDP
by: Yang, Xi, et al.
Published: (2026)
by: Yang, Xi, et al.
Published: (2026)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
Think Locally, Explain Globally: Graph-Guided LLM Investigations via Local Reasoning and Belief Propagation
by: Jha, Saurabh, et al.
Published: (2026)
by: Jha, Saurabh, et al.
Published: (2026)
Figaro on GPUs: Two Tables
by: Zivanovic, Dorde
Published: (2025)
by: Zivanovic, Dorde
Published: (2025)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
by: Sada, Mohammad Firas, et al.
Published: (2025)
by: Sada, Mohammad Firas, et al.
Published: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
by: Pan, Qilong, et al.
Published: (2025)
by: Pan, Qilong, et al.
Published: (2025)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
by: Le, Truong-Thanh, et al.
Published: (2026)
by: Le, Truong-Thanh, et al.
Published: (2026)
SUperman: Efficient Permanent Computation on GPUs
by: Elbek, Deniz, et al.
Published: (2025)
by: Elbek, Deniz, et al.
Published: (2025)
Control Flow Management in Modern GPUs
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
BDDC Preconditioning on GPUs for Cardiac Simulations
by: Goebel, Fritz, et al.
Published: (2024)
by: Goebel, Fritz, et al.
Published: (2024)
Input-Dependent Power Usage in GPUs
by: Gregersen, Theo, et al.
Published: (2024)
by: Gregersen, Theo, et al.
Published: (2024)
Modeling and Optimization of Control Problems on GPUs
by: Montoison, Alexis, et al.
Published: (2025)
by: Montoison, Alexis, et al.
Published: (2025)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023)
by: Kandel, Narayan, et al.
Published: (2023)
Similar Items
-
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025) -
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025) -
Queue management for slo-oriented large language model serving
by: Patke, Archit, et al.
Published: (2024) -
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
by: Qiu, Haoran, et al.
Published: (2024) -
Security Testbed for Preempting Attacks against Supercomputing Infrastructure
by: Cao, Phuong, et al.
Published: (2024)