On the Impact of Calibration Data in Post-training Quantization and Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Williams, Miles, Aletras, Nikolaos
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915004765175808
author Williams, Miles
Aletras, Nikolaos
author_facet Williams, Miles
Aletras, Nikolaos
contents Quantization and pruning form the foundation of compression for neural networks, enabling efficient inference for large language models (LLMs). Recently, various quantization and pruning techniques have demonstrated remarkable performance in a post-training setting. They rely upon calibration data, a small set of unlabeled examples that are used to generate layer activations. However, no prior work has systematically investigated how the calibration data impacts the effectiveness of model compression methods. In this paper, we present the first extensive empirical study on the effect of calibration data upon LLM performance. We trial a variety of quantization and pruning methods, datasets, tasks, and models. Surprisingly, we find substantial variations in downstream task performance, contrasting existing work that suggests a greater level of robustness to the calibration data. Finally, we make a series of recommendations for the effective use of calibration data in LLM quantization and pruning.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09755
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Impact of Calibration Data in Post-training Quantization and Pruning
Williams, Miles
Aletras, Nikolaos
Computation and Language
Quantization and pruning form the foundation of compression for neural networks, enabling efficient inference for large language models (LLMs). Recently, various quantization and pruning techniques have demonstrated remarkable performance in a post-training setting. They rely upon calibration data, a small set of unlabeled examples that are used to generate layer activations. However, no prior work has systematically investigated how the calibration data impacts the effectiveness of model compression methods. In this paper, we present the first extensive empirical study on the effect of calibration data upon LLM performance. We trial a variety of quantization and pruning methods, datasets, tasks, and models. Surprisingly, we find substantial variations in downstream task performance, contrasting existing work that suggests a greater level of robustness to the calibration data. Finally, we make a series of recommendations for the effective use of calibration data in LLM quantization and pruning.
title On the Impact of Calibration Data in Post-training Quantization and Pruning
topic Computation and Language
url https://arxiv.org/abs/2311.09755