Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xin, Meng, Priyadarshi, Sweta, Xin, Jingyu, Kartal, Bilal, Vavre, Aditya, Thekkumpate, Asma Kuriparambil, Chen, Zijia, Mahabaleshwarkar, Ameya Sunil, Shahaf, Ido, Bercovich, Akhiad, Patel, Kinjal, Velury, Suguna Varshini, Luo, Chenjie, Cheng, Zhiyu, Chen, Jenny, Yu, Chen-Han, Ping, Wei, Rybakov, Oleg, Tajbakhsh, Nima, Olabiyi, Oluwatobi, Stosic, Dusan, Wu, Di, Han, Song, Chung, Eric, Sreenivas, Sharath Turuvekere, Catanzaro, Bryan, Suhara, Yoshi, Blankevoort, Tijmen, Mao, Huizi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910038619062272
author Xin, Meng
Priyadarshi, Sweta
Xin, Jingyu
Kartal, Bilal
Vavre, Aditya
Thekkumpate, Asma Kuriparambil
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Shahaf, Ido
Bercovich, Akhiad
Patel, Kinjal
Velury, Suguna Varshini
Luo, Chenjie
Cheng, Zhiyu
Chen, Jenny
Yu, Chen-Han
Ping, Wei
Rybakov, Oleg
Tajbakhsh, Nima
Olabiyi, Oluwatobi
Stosic, Dusan
Wu, Di
Han, Song
Chung, Eric
Sreenivas, Sharath Turuvekere
Catanzaro, Bryan
Suhara, Yoshi
Blankevoort, Tijmen
Mao, Huizi
author_facet Xin, Meng
Priyadarshi, Sweta
Xin, Jingyu
Kartal, Bilal
Vavre, Aditya
Thekkumpate, Asma Kuriparambil
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Shahaf, Ido
Bercovich, Akhiad
Patel, Kinjal
Velury, Suguna Varshini
Luo, Chenjie
Cheng, Zhiyu
Chen, Jenny
Yu, Chen-Han
Ping, Wei
Rybakov, Oleg
Tajbakhsh, Nima
Olabiyi, Oluwatobi
Stosic, Dusan
Wu, Di
Han, Song
Chung, Eric
Sreenivas, Sharath Turuvekere
Catanzaro, Bryan
Suhara, Yoshi
Blankevoort, Tijmen
Mao, Huizi
contents This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying distillation to quantized models is not a new idea, we observe key advantages of QAD for today's LLMs: 1. It shows remarkable effectiveness and stability for models trained through multi-stage post-training pipelines, including supervised fine-tuning (SFT), reinforcement learning (RL), and model merging, where traditional quantization-aware training (QAT) suffers from engineering complexity and training instability; 2. It is robust to data quality and coverage, enabling accuracy recovery without full training data. We evaluate QAD across multiple post-trained models including AceReason Nemotron, Nemotron 3 Nano, Nemotron Nano V2, Nemotron Nano V2 VL (VLM), and Llama Nemotron Super v1, showing consistent recovery to near-BF16 accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20088
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
Xin, Meng
Priyadarshi, Sweta
Xin, Jingyu
Kartal, Bilal
Vavre, Aditya
Thekkumpate, Asma Kuriparambil
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Shahaf, Ido
Bercovich, Akhiad
Patel, Kinjal
Velury, Suguna Varshini
Luo, Chenjie
Cheng, Zhiyu
Chen, Jenny
Yu, Chen-Han
Ping, Wei
Rybakov, Oleg
Tajbakhsh, Nima
Olabiyi, Oluwatobi
Stosic, Dusan
Wu, Di
Han, Song
Chung, Eric
Sreenivas, Sharath Turuvekere
Catanzaro, Bryan
Suhara, Yoshi
Blankevoort, Tijmen
Mao, Huizi
Machine Learning
This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying distillation to quantized models is not a new idea, we observe key advantages of QAD for today's LLMs: 1. It shows remarkable effectiveness and stability for models trained through multi-stage post-training pipelines, including supervised fine-tuning (SFT), reinforcement learning (RL), and model merging, where traditional quantization-aware training (QAT) suffers from engineering complexity and training instability; 2. It is robust to data quality and coverage, enabling accuracy recovery without full training data. We evaluate QAD across multiple post-trained models including AceReason Nemotron, Nemotron 3 Nano, Nemotron Nano V2, Nemotron Nano V2 VL (VLM), and Llama Nemotron Super v1, showing consistent recovery to near-BF16 accuracy.
title Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
topic Machine Learning
url https://arxiv.org/abs/2601.20088