UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: von Rad, Jonathan, Cao, Yong, Geiger, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911712403259392
author von Rad, Jonathan
Cao, Yong
Geiger, Andreas
author_facet von Rad, Jonathan
Cao, Yong
Geiger, Andreas
contents Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primarily on knowledge-centric benchmarks. Thus, we introduce UniComp, a unified evaluation framework for comparing pruning, quantization, and knowledge distillation. UniComp evaluates compressed models along three dimensions: performance, reliability, and efficiency, using a diverse set of capability- and safety-oriented benchmarks together with a hardware-aware efficiency analysis. Through evaluation of six compression techniques across 40 datasets, we observe (i) a consistent knowledge bias, where factual recall is largely preserved while multi-step reasoning, multilingual, and instruction-following capabilities degrade; (ii) a decoupling between performance and reliability, indicating that retained performance does not consistently imply preserved reliability; and (iii) that task-specific calibration can yield up to 50% relative improvement of reasoning performance in pruned models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09130
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
von Rad, Jonathan
Cao, Yong
Geiger, Andreas
Machine Learning
Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primarily on knowledge-centric benchmarks. Thus, we introduce UniComp, a unified evaluation framework for comparing pruning, quantization, and knowledge distillation. UniComp evaluates compressed models along three dimensions: performance, reliability, and efficiency, using a diverse set of capability- and safety-oriented benchmarks together with a hardware-aware efficiency analysis. Through evaluation of six compression techniques across 40 datasets, we observe (i) a consistent knowledge bias, where factual recall is largely preserved while multi-step reasoning, multilingual, and instruction-following capabilities degrade; (ii) a decoupling between performance and reliability, indicating that retained performance does not consistently imply preserved reliability; and (iii) that task-specific calibration can yield up to 50% relative improvement of reasoning performance in pruned models.
title UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
topic Machine Learning
url https://arxiv.org/abs/2602.09130