HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sukthanker, Rhea Sanjay, Zela, Arber, Staffler, Benedikt, Klein, Aaron, Purucker, Lennart, Franke, Joerg K. H., Hutter, Frank
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917826219999232
author Sukthanker, Rhea Sanjay
Zela, Arber
Staffler, Benedikt
Klein, Aaron
Purucker, Lennart
Franke, Joerg K. H.
Hutter, Frank
author_facet Sukthanker, Rhea Sanjay
Zela, Arber
Staffler, Benedikt
Klein, Aaron
Purucker, Lennart
Franke, Joerg K. H.
Hutter, Frank
contents The increasing size of language models necessitates a thorough analysis across multiple dimensions to assess trade-offs among crucial hardware metrics such as latency, energy consumption, GPU memory usage, and performance. Identifying optimal model configurations under specific hardware constraints is becoming essential but remains challenging due to the computational load of exhaustive training and evaluation on multiple devices. To address this, we introduce HW-GPT-Bench, a hardware-aware benchmark that utilizes surrogate predictions to approximate various hardware metrics across 13 devices of architectures in the GPT-2 family, with architectures containing up to 1.55B parameters. Our surrogates, via calibrated predictions and reliable uncertainty estimates, faithfully model the heteroscedastic noise inherent in the energy and latency measurements. To estimate perplexity, we employ weight-sharing techniques from Neural Architecture Search (NAS), inheriting pretrained weights from the largest GPT-2 model. Finally, we demonstrate the utility of HW-GPT-Bench by simulating optimization trajectories of various multi-objective optimization algorithms in just a few seconds.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10299
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
Sukthanker, Rhea Sanjay
Zela, Arber
Staffler, Benedikt
Klein, Aaron
Purucker, Lennart
Franke, Joerg K. H.
Hutter, Frank
Machine Learning
Artificial Intelligence
The increasing size of language models necessitates a thorough analysis across multiple dimensions to assess trade-offs among crucial hardware metrics such as latency, energy consumption, GPU memory usage, and performance. Identifying optimal model configurations under specific hardware constraints is becoming essential but remains challenging due to the computational load of exhaustive training and evaluation on multiple devices. To address this, we introduce HW-GPT-Bench, a hardware-aware benchmark that utilizes surrogate predictions to approximate various hardware metrics across 13 devices of architectures in the GPT-2 family, with architectures containing up to 1.55B parameters. Our surrogates, via calibrated predictions and reliable uncertainty estimates, faithfully model the heteroscedastic noise inherent in the energy and latency measurements. To estimate perplexity, we employ weight-sharing techniques from Neural Architecture Search (NAS), inheriting pretrained weights from the largest GPT-2 model. Finally, we demonstrate the utility of HW-GPT-Bench by simulating optimization trajectories of various multi-objective optimization algorithms in just a few seconds.
title HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.10299