Scaling Laws for Post Training Quantized Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Zifei, Lan, Alexander, Yazar, Wanzin, Webb, Tristan, Sharify, Sayeh, Wang, Xin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915050042687488
author Xu, Zifei
Lan, Alexander
Yazar, Wanzin
Webb, Tristan
Sharify, Sayeh
Wang, Xin
author_facet Xu, Zifei
Lan, Alexander
Yazar, Wanzin
Webb, Tristan
Sharify, Sayeh
Wang, Xin
contents Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling laws governing pre-training, the quality of LLMs after post-training compression remains highly unpredictable, often requiring case-by-case validation in practice. In this work, we attempted to close this gap for post-training weight quantization of LLMs by conducting a systematic empirical study on multiple LLM families quantized to numerous low-precision tensor data types using popular weight quantization techniques. We identified key scaling factors pertaining to characteristics of the local loss landscape, based on which the performance of quantized LLMs can be reasonably well predicted by a statistical model.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12119
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Laws for Post Training Quantized Large Language Models
Xu, Zifei
Lan, Alexander
Yazar, Wanzin
Webb, Tristan
Sharify, Sayeh
Wang, Xin
Machine Learning
Computation and Language
Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling laws governing pre-training, the quality of LLMs after post-training compression remains highly unpredictable, often requiring case-by-case validation in practice. In this work, we attempted to close this gap for post-training weight quantization of LLMs by conducting a systematic empirical study on multiple LLM families quantized to numerous low-precision tensor data types using popular weight quantization techniques. We identified key scaling factors pertaining to characteristics of the local loss landscape, based on which the performance of quantized LLMs can be reasonably well predicted by a statistical model.
title Scaling Laws for Post Training Quantized Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.12119