Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qianli, Wang, Mingyang, Feldhus, Nils, Ostermann, Simon, Cao, Yuan, Schütze, Hinrich, Möller, Sebastian, Schmitt, Vera
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909000333787136
author Wang, Qianli
Wang, Mingyang
Feldhus, Nils
Ostermann, Simon
Cao, Yuan
Schütze, Hinrich
Möller, Sebastian
Schmitt, Vera
author_facet Wang, Qianli
Wang, Mingyang
Feldhus, Nils
Ostermann, Simon
Cao, Yuan
Schütze, Hinrich
Möller, Sebastian
Schmitt, Vera
contents Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilities have been extensively studied, one critical area remains underexplored: factual knowledge recall (FKR), the process by which LLMs access stored knowledge. To this end, we conduct comprehensive experiments using three common quantization techniques at distinct bit widths, in conjunction with interpretability-driven analyses on two tasks, knowledge memorization and latent multi-hop reasoning. We show that quantization typically results in information loss within LLMs, consequently diminishing their capacity for FKR. This effect is particularly amplified in smaller models within the same architectural families. However, models quantized at reduced bit precision do not consistently exhibit inferior performance and occasionally quantization may even enhance model FKR. We find that BitSandBytes demonstrates highest preservation of the original full-precision model's FKR. Despite variability across models and methods, quantization causes modest performance degradation and remains an effective compression strategy.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
Wang, Qianli
Wang, Mingyang
Feldhus, Nils
Ostermann, Simon
Cao, Yuan
Schütze, Hinrich
Möller, Sebastian
Schmitt, Vera
Computation and Language
Machine Learning
Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilities have been extensively studied, one critical area remains underexplored: factual knowledge recall (FKR), the process by which LLMs access stored knowledge. To this end, we conduct comprehensive experiments using three common quantization techniques at distinct bit widths, in conjunction with interpretability-driven analyses on two tasks, knowledge memorization and latent multi-hop reasoning. We show that quantization typically results in information loss within LLMs, consequently diminishing their capacity for FKR. This effect is particularly amplified in smaller models within the same architectural families. However, models quantized at reduced bit precision do not consistently exhibit inferior performance and occasionally quantization may even enhance model FKR. We find that BitSandBytes demonstrates highest preservation of the original full-precision model's FKR. Despite variability across models and methods, quantization causes modest performance degradation and remains an effective compression strategy.
title Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.13963