Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Jianwei, Yin, Hang, Deng, Peng, Almeida, Aline, Zhou, Shunfan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
di: Luo, Weile, et al.
Pubblicazione: (2025)
di: Luo, Weile, et al.
Pubblicazione: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
di: Cui, Shengkun, et al.
Pubblicazione: (2025)
di: Cui, Shengkun, et al.
Pubblicazione: (2025)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
di: Afzal, Ayesha, et al.
Pubblicazione: (2025)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
di: Sun, Bowen, et al.
Pubblicazione: (2026)
di: Sun, Bowen, et al.
Pubblicazione: (2026)
Can Large Language Models Predict Parallel Code Performance?
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
di: Sharma, Chandan, et al.
Pubblicazione: (2025)
di: Sharma, Chandan, et al.
Pubblicazione: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
di: Li, Junjie, et al.
Pubblicazione: (2024)
di: Li, Junjie, et al.
Pubblicazione: (2024)
How to Rent GPUs on a Budget
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
di: Owen, Herbert, et al.
Pubblicazione: (2024)
di: Owen, Herbert, et al.
Pubblicazione: (2024)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
di: Kulkarni, Apurv Deepak, et al.
Pubblicazione: (2025)
di: Kulkarni, Apurv Deepak, et al.
Pubblicazione: (2025)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
di: Renney, Harri, et al.
Pubblicazione: (2026)
di: Renney, Harri, et al.
Pubblicazione: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
di: Lin, Wei-Chen, et al.
Pubblicazione: (2024)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
di: León-Vega, Luis G., et al.
Pubblicazione: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
di: Rose, Martin, et al.
Pubblicazione: (2025)
di: Rose, Martin, et al.
Pubblicazione: (2025)
AGOCS -- Accurate Google Cloud Simulator Framework
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
Standardized Methods and Recommendations for Green Federated Learning
di: Tapp, Austin, et al.
Pubblicazione: (2026)
di: Tapp, Austin, et al.
Pubblicazione: (2026)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
di: Pinnock, Alyssa, et al.
Pubblicazione: (2025)
di: Pinnock, Alyssa, et al.
Pubblicazione: (2025)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
di: Barron, Ryan, et al.
Pubblicazione: (2024)
di: Barron, Ryan, et al.
Pubblicazione: (2024)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
di: Suo, Jiashun, et al.
Pubblicazione: (2025)
di: Suo, Jiashun, et al.
Pubblicazione: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)
di: Shen, Zixu, et al.
Pubblicazione: (2025)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
di: Bolet, Gregory, et al.
Pubblicazione: (2025)
When AI Bends Metal: AI-Assisted Optimization of Design Parameters in Sheet Metal Forming
di: Tarraf, Ahmad, et al.
Pubblicazione: (2025)
di: Tarraf, Ahmad, et al.
Pubblicazione: (2025)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
di: Dagli, Ismet, et al.
Pubblicazione: (2023)
di: Dagli, Ismet, et al.
Pubblicazione: (2023)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
di: Lurati, Milo, et al.
Pubblicazione: (2024)
di: Lurati, Milo, et al.
Pubblicazione: (2024)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
Democratizing AI: A Comparative Study in Deep Learning Efficiency and Future Trends in Computational Processing
di: Amin, Lisan Al, et al.
Pubblicazione: (2026)
di: Amin, Lisan Al, et al.
Pubblicazione: (2026)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
di: Panigrahy, Deepak, et al.
Pubblicazione: (2026)
di: Panigrahy, Deepak, et al.
Pubblicazione: (2026)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
di: Wang, Zhe, et al.
Pubblicazione: (2024)
di: Wang, Zhe, et al.
Pubblicazione: (2024)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
di: Liu, Hang, et al.
Pubblicazione: (2025)
di: Liu, Hang, et al.
Pubblicazione: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
di: Chakraborty, Abhinaba, et al.
Pubblicazione: (2025)
di: Chakraborty, Abhinaba, et al.
Pubblicazione: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
di: Tharwani, Jay, et al.
Pubblicazione: (2024)
di: Tharwani, Jay, et al.
Pubblicazione: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
di: Andersson, Måns I., et al.
Pubblicazione: (2025)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
di: Tang, Xinsheng, et al.
Pubblicazione: (2026)
di: Tang, Xinsheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
di: Luo, Weile, et al.
Pubblicazione: (2025) -
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
di: Cui, Shengkun, et al.
Pubblicazione: (2025) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
di: Afzal, Ayesha, et al.
Pubblicazione: (2025) -
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
di: Sun, Bowen, et al.
Pubblicazione: (2026) -
Can Large Language Models Predict Parallel Code Performance?
di: Bolet, Gregory, et al.
Pubblicazione: (2025)