Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Varshney, Ayush K., Vandikas, Konstantinos, Girdzijauskas, Šarūnas, Orucu, Adam, Feljan, Aneta Vulgarakis
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910252546392064
author Varshney, Ayush K.
Vandikas, Konstantinos
Girdzijauskas, Šarūnas
Orucu, Adam
Feljan, Aneta Vulgarakis
author_facet Varshney, Ayush K.
Vandikas, Konstantinos
Girdzijauskas, Šarūnas
Orucu, Adam
Feljan, Aneta Vulgarakis
contents Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss. Quantization-Aware Training (QAT) has emerged as a leading compression approach; however, existing mixed-precision methods typically operate at coarse layer- or channel-level granularity. These methods often rely on heuristic or search-based bit-allocation strategies, which may overlook fine-grained variability at the neuron level. We propose Neuron-Level Mixed-Precision QAT (NMP-QAT), where each neuron independently learns its own discrete precision during training. Starting from low-bit precision, NMP-QAT expands bit-width only when training signals demand it, via differentiable surrogates and straight-through estimators, while preserving a fully discrete inference graph. This adaptability extends to both weights and activations, reducing memory movement. Evaluated on telecom and non-telecom datasets across MLP and tabular foundation model architectures, NMP-QAT achieves superior compression-accuracy trade-offs over mixed-precision QAT baselines, making it well-suited for Green AI deployments at the network edge.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25054
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
Varshney, Ayush K.
Vandikas, Konstantinos
Girdzijauskas, Šarūnas
Orucu, Adam
Feljan, Aneta Vulgarakis
Machine Learning
Artificial Intelligence
Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss. Quantization-Aware Training (QAT) has emerged as a leading compression approach; however, existing mixed-precision methods typically operate at coarse layer- or channel-level granularity. These methods often rely on heuristic or search-based bit-allocation strategies, which may overlook fine-grained variability at the neuron level. We propose Neuron-Level Mixed-Precision QAT (NMP-QAT), where each neuron independently learns its own discrete precision during training. Starting from low-bit precision, NMP-QAT expands bit-width only when training signals demand it, via differentiable surrogates and straight-through estimators, while preserving a fully discrete inference graph. This adaptability extends to both weights and activations, reducing memory movement. Evaluated on telecom and non-telecom datasets across MLP and tabular foundation model architectures, NMP-QAT achieves superior compression-accuracy trade-offs over mixed-precision QAT baselines, making it well-suited for Green AI deployments at the network edge.
title Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.25054