Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sada, Mohammad Firas, Graham, John J., Khoda, Elham E, Tatineni, Mahidhar, Mishin, Dmitry, Gupta, Rajesh K., Wagner, Rick, Smarr, Larry, DeFanti, Thomas A., Würthwein, Frank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914120680341504
author Sada, Mohammad Firas
Graham, John J.
Khoda, Elham E
Tatineni, Mahidhar
Mishin, Dmitry
Gupta, Rajesh K.
Wagner, Rick
Smarr, Larry
DeFanti, Thomas A.
Würthwein, Frank
author_facet Sada, Mohammad Firas
Graham, John J.
Khoda, Elham E
Tatineni, Mahidhar
Mishin, Dmitry
Gupta, Rajesh K.
Wagner, Rick
Smarr, Larry
DeFanti, Thomas A.
Würthwein, Frank
contents This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP).
format Preprint
id arxiv_https___arxiv_org_abs_2507_00418
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
Sada, Mohammad Firas
Graham, John J.
Khoda, Elham E
Tatineni, Mahidhar
Mishin, Dmitry
Gupta, Rajesh K.
Wagner, Rick
Smarr, Larry
DeFanti, Thomas A.
Würthwein, Frank
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
68M20, 68T50
C.4; I.2.7; D.4.8
This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP).
title Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
68M20, 68T50
C.4; I.2.7; D.4.8
url https://arxiv.org/abs/2507.00418