Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Yingyu, Long, Jiangxuan, Shi, Zhenmei, Song, Zhao, Zhou, Yufa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916631566876672
author Liang, Yingyu
Long, Jiangxuan
Shi, Zhenmei
Song, Zhao
Zhou, Yufa
author_facet Liang, Yingyu
Long, Jiangxuan
Shi, Zhenmei
Song, Zhao
Zhou, Yufa
contents Large Language Models (LLMs) have shown immense potential in enhancing various aspects of our daily lives, from conversational AI to search and AI assistants. However, their growing capabilities come at the cost of extremely large model sizes, making deployment on edge devices challenging due to memory and computational constraints. This paper introduces a novel approach to LLM weight pruning that directly optimizes for approximating the attention matrix, a core component of transformer architectures. Unlike existing methods that focus on linear approximations, our approach accounts for the non-linear nature of the Softmax attention mechanism. We provide theoretical guarantees for the convergence of our Gradient Descent-based optimization method to a near-optimal pruning mask solution. Our empirical results demonstrate the effectiveness of our non-linear pruning approach in maintaining model performance while significantly reducing computational costs, which is beyond the current state-of-the-art methods, i.e., SparseGPT and Wanda, by a large margin. This work establishes a new theoretical foundation for pruning algorithm design in LLMs, potentially paving the way for more efficient LLM inference on resource-constrained devices.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11261
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
Liang, Yingyu
Long, Jiangxuan
Shi, Zhenmei
Song, Zhao
Zhou, Yufa
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have shown immense potential in enhancing various aspects of our daily lives, from conversational AI to search and AI assistants. However, their growing capabilities come at the cost of extremely large model sizes, making deployment on edge devices challenging due to memory and computational constraints. This paper introduces a novel approach to LLM weight pruning that directly optimizes for approximating the attention matrix, a core component of transformer architectures. Unlike existing methods that focus on linear approximations, our approach accounts for the non-linear nature of the Softmax attention mechanism. We provide theoretical guarantees for the convergence of our Gradient Descent-based optimization method to a near-optimal pruning mask solution. Our empirical results demonstrate the effectiveness of our non-linear pruning approach in maintaining model performance while significantly reducing computational costs, which is beyond the current state-of-the-art methods, i.e., SparseGPT and Wanda, by a large margin. This work establishes a new theoretical foundation for pruning algorithm design in LLMs, potentially paving the way for more efficient LLM inference on resource-constrained devices.
title Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.11261