LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hoffmann, David, Budhathoki, Kailash, Kleindessner, Matthaeus
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916499183108096
author Hoffmann, David
Budhathoki, Kailash
Kleindessner, Matthaeus
author_facet Hoffmann, David
Budhathoki, Kailash
Kleindessner, Matthaeus
contents The evolving capabilities of large language models are accompanied by growing sizes and deployment costs, necessitating effective inference optimisation techniques. We propose a novel pruning method utilising centrality measures from graph theory, reducing both the computational requirements and the memory footprint of these models. Specifically, we devise a method for creating a weighted directed acyclical graph representation of multilayer perceptrons to which we apply a modified version of the weighted PageRank centrality measure to compute node importance scores. In combination with uniform pruning this leads to structured sparsity. We call this pruning method MLPRank. Furthermore we introduce an extension to decoder-only transformer models and call it LLMRank. For both variants we demonstrate a strong performance. With MLPRank on average leading to 6.09 % higher accuracy retention than three popular baselines and 13.42 % with LLMRank compared to two popular baselines. Code is available at https://github.com/amazon-science/llm-rank-pruning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13299
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
Hoffmann, David
Budhathoki, Kailash
Kleindessner, Matthaeus
Machine Learning
Artificial Intelligence
The evolving capabilities of large language models are accompanied by growing sizes and deployment costs, necessitating effective inference optimisation techniques. We propose a novel pruning method utilising centrality measures from graph theory, reducing both the computational requirements and the memory footprint of these models. Specifically, we devise a method for creating a weighted directed acyclical graph representation of multilayer perceptrons to which we apply a modified version of the weighted PageRank centrality measure to compute node importance scores. In combination with uniform pruning this leads to structured sparsity. We call this pruning method MLPRank. Furthermore we introduce an extension to decoder-only transformer models and call it LLMRank. For both variants we demonstrate a strong performance. With MLPRank on average leading to 6.09 % higher accuracy retention than three popular baselines and 13.42 % with LLMRank compared to two popular baselines. Code is available at https://github.com/amazon-science/llm-rank-pruning.
title LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.13299