S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chavan, Arnav, Lele, Nahush, Bamba, Udbhav, Dayal, Sankalp, Raghunathan, Aditi, Gupta, Deepak
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911449899597824
author Chavan, Arnav
Lele, Nahush
Bamba, Udbhav
Dayal, Sankalp
Raghunathan, Aditi
Gupta, Deepak
author_facet Chavan, Arnav
Lele, Nahush
Bamba, Udbhav
Dayal, Sankalp
Raghunathan, Aditi
Gupta, Deepak
contents Activation outliers in large-scale transformer models pose a fundamental challenge to model quantization, creating excessively large ranges that cause severe accuracy drops during quantization. We empirically observe that outlier severity intensifies with pre-training scale (e.g., progressing from CLIP to the more extensively trained SigLIP and SigLIP2). Through theoretical analysis as well as empirical correlation studies, we establish the direct link between these activation outliers and dominant singular values of the weights. Building on this insight, we propose Selective Spectral Decay ($S^2D$), a geometrically-principled conditioning method that surgically regularizes only the weight components corresponding to the largest singular values during fine-tuning. Through extensive experiments, we demonstrate that $S^2D$ significantly reduces activation outliers and produces well-conditioned representations that are inherently quantization-friendly. Models trained with $S^2D$ achieve up to 7% improved PTQ accuracy on ImageNet under W4A4 quantization and 4% gains when combined with QAT. These improvements also generalize across downstream tasks and vision-language models, enabling the scaling of increasingly large and rigorously trained models without sacrificing deployment efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14432
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
Chavan, Arnav
Lele, Nahush
Bamba, Udbhav
Dayal, Sankalp
Raghunathan, Aditi
Gupta, Deepak
Machine Learning
Artificial Intelligence
Activation outliers in large-scale transformer models pose a fundamental challenge to model quantization, creating excessively large ranges that cause severe accuracy drops during quantization. We empirically observe that outlier severity intensifies with pre-training scale (e.g., progressing from CLIP to the more extensively trained SigLIP and SigLIP2). Through theoretical analysis as well as empirical correlation studies, we establish the direct link between these activation outliers and dominant singular values of the weights. Building on this insight, we propose Selective Spectral Decay ($S^2D$), a geometrically-principled conditioning method that surgically regularizes only the weight components corresponding to the largest singular values during fine-tuning. Through extensive experiments, we demonstrate that $S^2D$ significantly reduces activation outliers and produces well-conditioned representations that are inherently quantization-friendly. Models trained with $S^2D$ achieve up to 7% improved PTQ accuracy on ImageNet under W4A4 quantization and 4% gains when combined with QAT. These improvements also generalize across downstream tasks and vision-language models, enabling the scaling of increasingly large and rigorously trained models without sacrificing deployment efficiency.
title S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.14432