LLM Compression: How Far Can We Go in Balancing Size and Performance?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sk, Sahil, Dhal, Debasish, Khosla, Sonal, Shahid, Sk, Shekhar, Sambit, Dhaka, Akash, Parida, Shantipriya, Prasad, Dilip K., Bojar, Ondřej
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916901341364224
author Sk, Sahil
Dhal, Debasish
Khosla, Sonal
Shahid, Sk
Shekhar, Sambit
Dhaka, Akash
Parida, Shantipriya
Prasad, Dilip K.
Bojar, Ondřej
author_facet Sk, Sahil
Dhal, Debasish
Khosla, Sonal
Shahid, Sk
Shekhar, Sambit
Dhaka, Akash
Parida, Shantipriya
Prasad, Dilip K.
Bojar, Ondřej
contents Quantization is an essential and popular technique for improving the accessibility of large language models (LLMs) by reducing memory usage and computational costs while maintaining performance. In this study, we apply 4-bit Group Scaling Quantization (GSQ) and Generative Pretrained Transformer Quantization (GPTQ) to LLaMA 1B, Qwen 0.5B, and PHI 1.5B, evaluating their impact across multiple NLP tasks. We benchmark these models on MS MARCO (Information Retrieval), BoolQ (Boolean Question Answering), and GSM8K (Mathematical Reasoning) datasets, assessing both accuracy and efficiency across various tasks. The study measures the trade-offs between model compression and task performance, analyzing key evaluation metrics, namely accuracy, inference latency, and throughput (total output tokens generated per second), providing insights into the suitability of low-bit quantization for real-world deployment. Using the results, users can then make suitable decisions based on the specifications that need to be met. We discuss the pros and cons of GSQ and GPTQ techniques on models of different sizes, which also serve as a benchmark for future experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Compression: How Far Can We Go in Balancing Size and Performance?
Sk, Sahil
Dhal, Debasish
Khosla, Sonal
Shahid, Sk
Shekhar, Sambit
Dhaka, Akash
Parida, Shantipriya
Prasad, Dilip K.
Bojar, Ondřej
Computation and Language
Quantization is an essential and popular technique for improving the accessibility of large language models (LLMs) by reducing memory usage and computational costs while maintaining performance. In this study, we apply 4-bit Group Scaling Quantization (GSQ) and Generative Pretrained Transformer Quantization (GPTQ) to LLaMA 1B, Qwen 0.5B, and PHI 1.5B, evaluating their impact across multiple NLP tasks. We benchmark these models on MS MARCO (Information Retrieval), BoolQ (Boolean Question Answering), and GSM8K (Mathematical Reasoning) datasets, assessing both accuracy and efficiency across various tasks. The study measures the trade-offs between model compression and task performance, analyzing key evaluation metrics, namely accuracy, inference latency, and throughput (total output tokens generated per second), providing insights into the suitability of low-bit quantization for real-world deployment. Using the results, users can then make suitable decisions based on the specifications that need to be met. We discuss the pros and cons of GSQ and GPTQ techniques on models of different sizes, which also serve as a benchmark for future experiments.
title LLM Compression: How Far Can We Go in Balancing Size and Performance?
topic Computation and Language
url https://arxiv.org/abs/2508.11318