Saved in:
Bibliographic Details
Main Authors: Cui, Wanyun, Wang, Qianle
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.02837
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929608686829568
author Cui, Wanyun
Wang, Qianle
author_facet Cui, Wanyun
Wang, Qianle
contents This paper reveals the phenomenon of parameter heterogeneity in large language models (LLMs). We find that a small subset of "cherry" parameters exhibit a disproportionately large influence on model performance, while the vast majority of parameters have minimal impact. This heterogeneity is found to be prevalent across different model families, scales, and types. Motivated by this observation, we propose CherryQ, a novel quantization method that unifies the optimization of mixed-precision parameters. CherryQ identifies and preserves the critical cherry parameters in high precision while aggressively quantizing the remaining parameters to low precision. Extensive experiments demonstrate the effectiveness of CherryQ. CherryQ outperforms existing quantization approaches in terms of perplexity and downstream task performance. Notably, our 3-bit quantized Vicuna-1.5 exhibits competitive performance compared to their 16-bit counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02837
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
Cui, Wanyun
Wang, Qianle
Computation and Language
This paper reveals the phenomenon of parameter heterogeneity in large language models (LLMs). We find that a small subset of "cherry" parameters exhibit a disproportionately large influence on model performance, while the vast majority of parameters have minimal impact. This heterogeneity is found to be prevalent across different model families, scales, and types. Motivated by this observation, we propose CherryQ, a novel quantization method that unifies the optimization of mixed-precision parameters. CherryQ identifies and preserves the critical cherry parameters in high precision while aggressively quantizing the remaining parameters to low precision. Extensive experiments demonstrate the effectiveness of CherryQ. CherryQ outperforms existing quantization approaches in terms of perplexity and downstream task performance. Notably, our 3-bit quantized Vicuna-1.5 exhibits competitive performance compared to their 16-bit counterparts.
title Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2404.02837