Saved in:
Bibliographic Details
Main Authors: Kim, Jiwoo, Lee, Joonhyung, Park, Gunho, Kim, Byeongwook, Kwon, Se Jung, Lee, Dongsoo, Lee, Youngjoo
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.01070
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911118112325632
author Kim, Jiwoo
Lee, Joonhyung
Park, Gunho
Kim, Byeongwook
Kwon, Se Jung
Lee, Dongsoo
Lee, Youngjoo
author_facet Kim, Jiwoo
Lee, Joonhyung
Park, Gunho
Kim, Byeongwook
Kwon, Se Jung
Lee, Dongsoo
Lee, Youngjoo
contents As large language models (LLMs) continue to scale, the high power consumption of AI accelerators in datacenters presents significant challenges, substantially increasing the total cost of ownership (TCO) for cloud service providers (CSPs) that provide LLM inference. In this work, we analyze the computational characteristics of LLM inference from a TCO perspective and present a generalizable framework to compare AI accelerators across diverse operational requirements. Using this model, we investigate key workload characteristics influencing TCO for AI accelerators from Intel (Gaudi 2 & 3) and NVIDIA (H100 & H200), especially thin GEMM utilization and FP8 quantization. In particular, as FP8 emerges as the baseline precision for next-generation LLMs, understanding how different architectures implement and benefit from low-precision computation is increasingly critical. Throughput on thin GEMMs has a greater impact on TCO than theoretical hardware peak throughput because the memory-bound decode phase is dominated by GEMV-like computations. We find that Gaudi HPUs achieve superior utilization on thin GEMMs compared to their counterparts, especially in FP8-quantized models. Our result underscores the importance of empirical, workload-level analysis in evaluating accelerator performance, rather than relying solely on theoretical hardware specifications. By studying the interaction between power consumption, quantization strategies, and hardware architecture, we provide insights to support informed deployment decisions and guide future accelerator designs aimed at improving the TCO of LLM inference workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Inquiry into Datacenter TCO for LLM Inference with FP8
Kim, Jiwoo
Lee, Joonhyung
Park, Gunho
Kim, Byeongwook
Kwon, Se Jung
Lee, Dongsoo
Lee, Youngjoo
Machine Learning
Performance
As large language models (LLMs) continue to scale, the high power consumption of AI accelerators in datacenters presents significant challenges, substantially increasing the total cost of ownership (TCO) for cloud service providers (CSPs) that provide LLM inference. In this work, we analyze the computational characteristics of LLM inference from a TCO perspective and present a generalizable framework to compare AI accelerators across diverse operational requirements. Using this model, we investigate key workload characteristics influencing TCO for AI accelerators from Intel (Gaudi 2 & 3) and NVIDIA (H100 & H200), especially thin GEMM utilization and FP8 quantization. In particular, as FP8 emerges as the baseline precision for next-generation LLMs, understanding how different architectures implement and benefit from low-precision computation is increasingly critical. Throughput on thin GEMMs has a greater impact on TCO than theoretical hardware peak throughput because the memory-bound decode phase is dominated by GEMV-like computations. We find that Gaudi HPUs achieve superior utilization on thin GEMMs compared to their counterparts, especially in FP8-quantized models. Our result underscores the importance of empirical, workload-level analysis in evaluating accelerator performance, rather than relying solely on theoretical hardware specifications. By studying the interaction between power consumption, quantization strategies, and hardware architecture, we provide insights to support informed deployment decisions and guide future accelerator designs aimed at improving the TCO of LLM inference workloads.
title An Inquiry into Datacenter TCO for LLM Inference with FP8
topic Machine Learning
Performance
url https://arxiv.org/abs/2502.01070