AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Sharma, Amit
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912420434280448
author Sharma, Amit
author_facet Sharma, Amit
contents The rapid growth of large-language models (LLMs) is driving a new wave of specialized hardware for inference. This paper presents the first workload-centric, cross-architectural performance study of commercial AI accelerators, spanning GPU-based chips, hybrid packages, and wafer-scale engines. We compare memory hierarchies, compute fabrics, and on-chip interconnects, and observe up to 3.7x performance variation across architectures as batch size and sequence length change. Four scaling techniques for trillion-parameter models are examined; expert parallelism offers an 8.4x parameter-to-compute advantage but incurs 2.1x higher latency variance than tensor parallelism. These findings provide quantitative guidance for matching workloads to accelerators and reveal architectural gaps that next-generation designs must address.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
Sharma, Amit
Hardware Architecture
Machine Learning
The rapid growth of large-language models (LLMs) is driving a new wave of specialized hardware for inference. This paper presents the first workload-centric, cross-architectural performance study of commercial AI accelerators, spanning GPU-based chips, hybrid packages, and wafer-scale engines. We compare memory hierarchies, compute fabrics, and on-chip interconnects, and observe up to 3.7x performance variation across architectures as batch size and sequence length change. Four scaling techniques for trillion-parameter models are examined; expert parallelism offers an 8.4x parameter-to-compute advantage but incurs 2.1x higher latency variance than tensor parallelism. These findings provide quantitative guidance for matching workloads to accelerators and reveal architectural gaps that next-generation designs must address.
title AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2506.00008