Enregistré dans:
Détails bibliographiques
Auteurs principaux: Recasens, Pol G., Zhu, Yue, Wang, Chen, Lee, Eun Kyung, Tardieu, Olivier, Youssef, Alaa, Torres, Jordi, Berral, Josep Ll.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2404.03353
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912526125498368
author Recasens, Pol G.
Zhu, Yue
Wang, Chen
Lee, Eun Kyung
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
author_facet Recasens, Pol G.
Zhu, Yue
Wang, Chen
Lee, Eun Kyung
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
contents Large language models (LLMs) have revolutionized the state-of-the-art of many different natural language processing tasks. Although serving LLMs is computationally and memory demanding, the rise of Small Language Models (SLMs) offers new opportunities for resource-constrained users, who now are able to serve small models with cutting-edge performance. In this paper, we present a set of experiments designed to benchmark SLM inference at performance and energy levels. Our analysis provides a new perspective in serving, highlighting that the small memory footprint of SLMs allows for reaching the Pareto-optimal throughput within the resource capacity of a single accelerator. In this regard, we present an initial set of findings demonstrating how model replication can effectively improve resource utilization for serving SLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Pareto Optimal Throughput in Small Language Model Serving
Recasens, Pol G.
Zhu, Yue
Wang, Chen
Lee, Eun Kyung
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
Computation and Language
Large language models (LLMs) have revolutionized the state-of-the-art of many different natural language processing tasks. Although serving LLMs is computationally and memory demanding, the rise of Small Language Models (SLMs) offers new opportunities for resource-constrained users, who now are able to serve small models with cutting-edge performance. In this paper, we present a set of experiments designed to benchmark SLM inference at performance and energy levels. Our analysis provides a new perspective in serving, highlighting that the small memory footprint of SLMs allows for reaching the Pareto-optimal throughput within the resource capacity of a single accelerator. In this regard, we present an initial set of findings demonstrating how model replication can effectively improve resource utilization for serving SLMs.
title Towards Pareto Optimal Throughput in Small Language Model Serving
topic Computation and Language
url https://arxiv.org/abs/2404.03353