Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910038619062272 |
|---|---|
| author | Xin, Meng Priyadarshi, Sweta Xin, Jingyu Kartal, Bilal Vavre, Aditya Thekkumpate, Asma Kuriparambil Chen, Zijia Mahabaleshwarkar, Ameya Sunil Shahaf, Ido Bercovich, Akhiad Patel, Kinjal Velury, Suguna Varshini Luo, Chenjie Cheng, Zhiyu Chen, Jenny Yu, Chen-Han Ping, Wei Rybakov, Oleg Tajbakhsh, Nima Olabiyi, Oluwatobi Stosic, Dusan Wu, Di Han, Song Chung, Eric Sreenivas, Sharath Turuvekere Catanzaro, Bryan Suhara, Yoshi Blankevoort, Tijmen Mao, Huizi |
| author_facet | Xin, Meng Priyadarshi, Sweta Xin, Jingyu Kartal, Bilal Vavre, Aditya Thekkumpate, Asma Kuriparambil Chen, Zijia Mahabaleshwarkar, Ameya Sunil Shahaf, Ido Bercovich, Akhiad Patel, Kinjal Velury, Suguna Varshini Luo, Chenjie Cheng, Zhiyu Chen, Jenny Yu, Chen-Han Ping, Wei Rybakov, Oleg Tajbakhsh, Nima Olabiyi, Oluwatobi Stosic, Dusan Wu, Di Han, Song Chung, Eric Sreenivas, Sharath Turuvekere Catanzaro, Bryan Suhara, Yoshi Blankevoort, Tijmen Mao, Huizi |
| contents | This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying distillation to quantized models is not a new idea, we observe key advantages of QAD for today's LLMs: 1. It shows remarkable effectiveness and stability for models trained through multi-stage post-training pipelines, including supervised fine-tuning (SFT), reinforcement learning (RL), and model merging, where traditional quantization-aware training (QAT) suffers from engineering complexity and training instability; 2. It is robust to data quality and coverage, enabling accuracy recovery without full training data. We evaluate QAD across multiple post-trained models including AceReason Nemotron, Nemotron 3 Nano, Nemotron Nano V2, Nemotron Nano V2 VL (VLM), and Llama Nemotron Super v1, showing consistent recovery to near-BF16 accuracy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_20088 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery Xin, Meng Priyadarshi, Sweta Xin, Jingyu Kartal, Bilal Vavre, Aditya Thekkumpate, Asma Kuriparambil Chen, Zijia Mahabaleshwarkar, Ameya Sunil Shahaf, Ido Bercovich, Akhiad Patel, Kinjal Velury, Suguna Varshini Luo, Chenjie Cheng, Zhiyu Chen, Jenny Yu, Chen-Han Ping, Wei Rybakov, Oleg Tajbakhsh, Nima Olabiyi, Oluwatobi Stosic, Dusan Wu, Di Han, Song Chung, Eric Sreenivas, Sharath Turuvekere Catanzaro, Bryan Suhara, Yoshi Blankevoort, Tijmen Mao, Huizi Machine Learning This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying distillation to quantized models is not a new idea, we observe key advantages of QAD for today's LLMs: 1. It shows remarkable effectiveness and stability for models trained through multi-stage post-training pipelines, including supervised fine-tuning (SFT), reinforcement learning (RL), and model merging, where traditional quantization-aware training (QAT) suffers from engineering complexity and training instability; 2. It is robust to data quality and coverage, enabling accuracy recovery without full training data. We evaluate QAD across multiple post-trained models including AceReason Nemotron, Nemotron 3 Nano, Nemotron Nano V2, Nemotron Nano V2 VL (VLM), and Llama Nemotron Super v1, showing consistent recovery to near-BF16 accuracy. |
| title | Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2601.20088 |