Performance of Confidential Computing GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ibarra, Antonio Martínez, Stephen, Julian James, Vidal, Aurora González, Jayaram, K. R., Gómez, Antonio Fernando Skarmeta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913852338208768
author Ibarra, Antonio Martínez
Stephen, Julian James
Vidal, Aurora González
Jayaram, K. R.
Gómez, Antonio Fernando Skarmeta
author_facet Ibarra, Antonio Martínez
Stephen, Julian James
Vidal, Aurora González
Jayaram, K. R.
Gómez, Antonio Fernando Skarmeta
contents This work examines latency, throughput, and other metrics when performing inference on confidential GPUs. We explore different traffic patterns and scheduling strategies using a single Virtual Machine with one NVIDIA H100 GPU, to perform relaxed batch inferences on multiple Large Language Models (LLMs), operating under the constraint of swapping models in and out of memory, which necessitates efficient control. The experiments simulate diverse real-world scenarios by varying parameters such as traffic load, traffic distribution patterns, scheduling strategies, and Service Level Agreement (SLA) requirements. The findings provide insights into the differences between confidential and non-confidential settings when performing inference in scenarios requiring active model swapping. Results indicate that in No-CC mode, relaxed batch inference with model swapping latency is 20-30% lower than in confidential mode. Additionally, SLA attainment is 15-20% higher in No-CC settings. Throughput in No-CC scenarios surpasses that of confidential mode by 45-70%, and GPU utilization is approximately 50% higher in No-CC environments. Overall, performance in the confidential setting is inferior to that in the No-CC scenario, primarily due to the additional encryption and decryption overhead required for loading models onto the GPU in confidential environments.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance of Confidential Computing GPUs
Ibarra, Antonio Martínez
Stephen, Julian James
Vidal, Aurora González
Jayaram, K. R.
Gómez, Antonio Fernando Skarmeta
Performance
This work examines latency, throughput, and other metrics when performing inference on confidential GPUs. We explore different traffic patterns and scheduling strategies using a single Virtual Machine with one NVIDIA H100 GPU, to perform relaxed batch inferences on multiple Large Language Models (LLMs), operating under the constraint of swapping models in and out of memory, which necessitates efficient control. The experiments simulate diverse real-world scenarios by varying parameters such as traffic load, traffic distribution patterns, scheduling strategies, and Service Level Agreement (SLA) requirements. The findings provide insights into the differences between confidential and non-confidential settings when performing inference in scenarios requiring active model swapping. Results indicate that in No-CC mode, relaxed batch inference with model swapping latency is 20-30% lower than in confidential mode. Additionally, SLA attainment is 15-20% higher in No-CC settings. Throughput in No-CC scenarios surpasses that of confidential mode by 45-70%, and GPU utilization is approximately 50% higher in No-CC environments. Overall, performance in the confidential setting is inferior to that in the No-CC scenario, primarily due to the additional encryption and decryption overhead required for loading models onto the GPU in confidential environments.
title Performance of Confidential Computing GPUs
topic Performance
url https://arxiv.org/abs/2505.16501