Compressing Large Language Models with Automated Sub-Network Search

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sukthanker, Rhea Sanjay, Staffler, Benedikt, Hutter, Frank, Klein, Aaron
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915139252387840
author Sukthanker, Rhea Sanjay
Staffler, Benedikt
Hutter, Frank
Klein, Aaron
author_facet Sukthanker, Rhea Sanjay
Staffler, Benedikt
Hutter, Frank
Klein, Aaron
contents Large Language Models (LLMs) demonstrate exceptional reasoning abilities, enabling strong generalization across diverse tasks such as commonsense reasoning and instruction following. However, as LLMs scale, inference costs become increasingly prohibitive, accumulating significantly over their life cycle. In this paper we consider model compression for LLMs to reduce model size while improving downstream task performance. We phrase this as a neural architecture search problem that automatically prunes structural components, such as attention heads, neurons, and layers by searching for the Pareto-optimal set of sub-networks balancing between performance and on-device latency. Compared to state-of-the-art structural pruning approaches and fine-tuned smaller sub-networks extracted from the pre-trained model, our method achieves upto 9.85% improvement on average on 11 diverse downstream tasks, while achieving up to 22% improvement of on-device latency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compressing Large Language Models with Automated Sub-Network Search
Sukthanker, Rhea Sanjay
Staffler, Benedikt
Hutter, Frank
Klein, Aaron
Computation and Language
Large Language Models (LLMs) demonstrate exceptional reasoning abilities, enabling strong generalization across diverse tasks such as commonsense reasoning and instruction following. However, as LLMs scale, inference costs become increasingly prohibitive, accumulating significantly over their life cycle. In this paper we consider model compression for LLMs to reduce model size while improving downstream task performance. We phrase this as a neural architecture search problem that automatically prunes structural components, such as attention heads, neurons, and layers by searching for the Pareto-optimal set of sub-networks balancing between performance and on-device latency. Compared to state-of-the-art structural pruning approaches and fine-tuned smaller sub-networks extracted from the pre-trained model, our method achieves upto 9.85% improvement on average on 11 diverse downstream tasks, while achieving up to 22% improvement of on-device latency.
title Compressing Large Language Models with Automated Sub-Network Search
topic Computation and Language
url https://arxiv.org/abs/2410.06479