Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914120680341504 |
|---|---|
| author | Sada, Mohammad Firas Graham, John J. Khoda, Elham E Tatineni, Mahidhar Mishin, Dmitry Gupta, Rajesh K. Wagner, Rick Smarr, Larry DeFanti, Thomas A. Würthwein, Frank |
| author_facet | Sada, Mohammad Firas Graham, John J. Khoda, Elham E Tatineni, Mahidhar Mishin, Dmitry Gupta, Rajesh K. Wagner, Rick Smarr, Larry DeFanti, Thomas A. Würthwein, Frank |
| contents | This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_00418 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs Sada, Mohammad Firas Graham, John J. Khoda, Elham E Tatineni, Mahidhar Mishin, Dmitry Gupta, Rajesh K. Wagner, Rick Smarr, Larry DeFanti, Thomas A. Würthwein, Frank Distributed, Parallel, and Cluster Computing Artificial Intelligence 68M20, 68T50 C.4; I.2.7; D.4.8 This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP). |
| title | Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence 68M20, 68T50 C.4; I.2.7; D.4.8 |
| url | https://arxiv.org/abs/2507.00418 |