Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hakim, Muhammad Alif Al, Wicaksono, Alfan Farizki, Koto, Fajri
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909993341550592
author Hakim, Muhammad Alif Al
Wicaksono, Alfan Farizki
Koto, Fajri
author_facet Hakim, Muhammad Alif Al
Wicaksono, Alfan Farizki
Koto, Fajri
contents Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantization and multilingual contexts, remain underexplored. In this work, we conduct a systematic study of how static and dynamic quantization methods impact fairness and safety across benchmarks measuring intrinsic and extrinsic bias and safety alignment. For fairness, we evaluate English, French, Dutch, Spanish, and Turkish; for safety, we focus on English, Korean, and Arabic. Our findings reveal that quantization consistently degrades fairness and safety, with dynamic methods demonstrating greater stability than static ones. Moreover, fairness degradation varies across languages, while safety deterioration is especially pronounced in non-English settings. To address these risks, we introduce Critical Weight Protection, a novel technique that identifies and preserves fairness- and safety-critical weights during quantization. This approach effectively mitigates bias and safety deterioration without costly retraining or alignment, maintaining trustworthiness while retaining efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12033
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
Hakim, Muhammad Alif Al
Wicaksono, Alfan Farizki
Koto, Fajri
Computation and Language
Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantization and multilingual contexts, remain underexplored. In this work, we conduct a systematic study of how static and dynamic quantization methods impact fairness and safety across benchmarks measuring intrinsic and extrinsic bias and safety alignment. For fairness, we evaluate English, French, Dutch, Spanish, and Turkish; for safety, we focus on English, Korean, and Arabic. Our findings reveal that quantization consistently degrades fairness and safety, with dynamic methods demonstrating greater stability than static ones. Moreover, fairness degradation varies across languages, while safety deterioration is especially pronounced in non-English settings. To address these risks, we introduce Critical Weight Protection, a novel technique that identifies and preserves fairness- and safety-critical weights during quantization. This approach effectively mitigates bias and safety deterioration without costly retraining or alignment, maintaining trustworthiness while retaining efficiency.
title Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
topic Computation and Language
url https://arxiv.org/abs/2601.12033