AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Song, Zichen, Wu, Yuxin, Huang, Sitan, Kang, Zhongfeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915004469477376
author Song, Zichen
Wu, Yuxin
Huang, Sitan
Kang, Zhongfeng
author_facet Song, Zichen
Wu, Yuxin
Huang, Sitan
Kang, Zhongfeng
contents The evaluation of layer importance in deep learning has been an active area of research, with significant implications for model optimization and interpretability. Recently, large language models (LLMs) have gained prominence across various domains, yet limited studies have explored the functional importance and performance contributions of individual layers within LLMs, especially from the perspective of activation distribution. In this work, we propose the Activation Variance-Sparsity Score (AVSS), a novel metric combining normalized activation variance and sparsity to assess each layer's contribution to model performance. By identifying and removing approximately the lowest 25% of layers based on AVSS, we achieve over 90% of original model performance across tasks such as question answering, language modeling, and sentiment classification, indicating that these layers may be non-essential. Our approach provides a systematic method for identifying less critical layers, contributing to efficient large language model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02117
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis
Song, Zichen
Wu, Yuxin
Huang, Sitan
Kang, Zhongfeng
Computation and Language
The evaluation of layer importance in deep learning has been an active area of research, with significant implications for model optimization and interpretability. Recently, large language models (LLMs) have gained prominence across various domains, yet limited studies have explored the functional importance and performance contributions of individual layers within LLMs, especially from the perspective of activation distribution. In this work, we propose the Activation Variance-Sparsity Score (AVSS), a novel metric combining normalized activation variance and sparsity to assess each layer's contribution to model performance. By identifying and removing approximately the lowest 25% of layers based on AVSS, we achieve over 90% of original model performance across tasks such as question answering, language modeling, and sentiment classification, indicating that these layers may be non-essential. Our approach provides a systematic method for identifying less critical layers, contributing to efficient large language model architectures.
title AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis
topic Computation and Language
url https://arxiv.org/abs/2411.02117