Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
Fuente:
arXiv
Saved in:
| Main Authors: | Xin, Meng, Priyadarshi, Sweta, Xin, Jingyu, Kartal, Bilal, Vavre, Aditya, Thekkumpate, Asma Kuriparambil, Chen, Zijia, Mahabaleshwarkar, Ameya Sunil, Shahaf, Ido, Bercovich, Akhiad, Patel, Kinjal, Velury, Suguna Varshini, Luo, Chenjie, Cheng, Zhiyu, Chen, Jenny, Yu, Chen-Han, Ping, Wei, Rybakov, Oleg, Tajbakhsh, Nima, Olabiyi, Oluwatobi, Stosic, Dusan, Wu, Di, Han, Song, Chung, Eric, Sreenivas, Sharath Turuvekere, Catanzaro, Bryan, Suhara, Yoshi, Blankevoort, Tijmen, Mao, Huizi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
by: Taghibakhshi, Ali, et al.
Published: (2026)
by: Taghibakhshi, Ali, et al.
Published: (2026)
BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
by: Venkatesh, Thejas, et al.
Published: (2026)
by: Venkatesh, Thejas, et al.
Published: (2026)
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
by: Taghibakhshi, Ali, et al.
Published: (2025)
by: Taghibakhshi, Ali, et al.
Published: (2025)
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
by: Taghibakhshi, Ali, et al.
Published: (2025)
by: Taghibakhshi, Ali, et al.
Published: (2025)
Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
by: Cook, Jack, et al.
Published: (2025)
by: Cook, Jack, et al.
Published: (2025)
Source Identification in Abstractive Summarization
by: Suhara, Yoshi, et al.
Published: (2024)
by: Suhara, Yoshi, et al.
Published: (2024)
Not All NVFP4 QAT Recipes Are Equal: How Architecture and Scale Shape Model Quality for Anomaly Segmentation
by: Du, Zijian, et al.
Published: (2026)
by: Du, Zijian, et al.
Published: (2026)
Changing Base Without Losing Pace: A GPU-Efficient Alternative to MatMul in DNNs
by: Ailon, Nir, et al.
Published: (2025)
by: Ailon, Nir, et al.
Published: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
by: Iso, Hayate, et al.
Published: (2022)
by: Iso, Hayate, et al.
Published: (2022)
Large Language Models are Inconsistent and Biased Evaluators
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
by: Kopiczko, Dawid J., et al.
Published: (2024)
by: Kopiczko, Dawid J., et al.
Published: (2024)
VeRA: Vector-based Random Matrix Adaptation
by: Kopiczko, Dawid J., et al.
Published: (2023)
by: Kopiczko, Dawid J., et al.
Published: (2023)
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
by: Chen, Yukang, et al.
Published: (2026)
by: Chen, Yukang, et al.
Published: (2026)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Hymba: A Hybrid-head Architecture for Small Language Models
by: Dong, Xin, et al.
Published: (2024)
by: Dong, Xin, et al.
Published: (2024)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
by: Kopiczko, Dawid J., et al.
Published: (2026)
by: Kopiczko, Dawid J., et al.
Published: (2026)
Your Large Language Models Are Leaving Fingerprints
by: McGovern, Hope, et al.
Published: (2024)
by: McGovern, Hope, et al.
Published: (2024)
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023)
by: Kuzmin, Andrey, et al.
Published: (2023)
FAAR: Format-Aware Adaptive Rounding for NVFP4
by: Li, Hanglin, et al.
Published: (2026)
by: Li, Hanglin, et al.
Published: (2026)
Dissecting Multifractal detrended cross-correlation analysis
by: Stosic, Borko, et al.
Published: (2024)
by: Stosic, Borko, et al.
Published: (2024)
Llama 3 Meets MoE: Efficient Upcycling
by: Vavre, Aditya, et al.
Published: (2024)
by: Vavre, Aditya, et al.
Published: (2024)
Discovering Governing Equations in the Presence of Uncertainty
by: Olabiyi, Ridwan, et al.
Published: (2025)
by: Olabiyi, Ridwan, et al.
Published: (2025)
CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
by: Uziel, Roy, et al.
Published: (2026)
by: Uziel, Roy, et al.
Published: (2026)
Blockchain for Cybersecurity: Securing Transactions, Data, and Identity
by: Suguna Balusamy
Published: (2025)
by: Suguna Balusamy
Published: (2025)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
by: Zou, Jiaxiang, et al.
Published: (2026)
by: Zou, Jiaxiang, et al.
Published: (2026)
Optimization of Wind Turbine Placement for Maximum Energy Output in Bangladesh
by: Winner, Olabiyi
Published: (2026)
by: Winner, Olabiyi
Published: (2026)
Policy and Investment Strategies for Expanding Wind Energy Projects in Bangladesh
by: Winner, Olabiyi
Published: (2026)
by: Winner, Olabiyi
Published: (2026)
Quasiparticle Variational Quantum Eigensolver
by: Velury, Saavanth, et al.
Published: (2025)
by: Velury, Saavanth, et al.
Published: (2025)
Thermodynamics of data
by: Stosic, Borko
Published: (2025)
by: Stosic, Borko
Published: (2025)
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
by: Bercovich, Ivan
Published: (2026)
by: Bercovich, Ivan
Published: (2026)
Generalized knots-quivers correspondence
by: Stošić, Marko
Published: (2024)
by: Stošić, Marko
Published: (2024)
Multiplication Operator Semigroups on Banach lattice valued continuous function spaces
by: Olabiyi, Tobi David
Published: (2025)
by: Olabiyi, Tobi David
Published: (2025)
Regioselective Ring Opening of Epoxides with Amines Using Silica-bonded S-sulfonic Acid under Solvent-free Conditions
by: Mahmood Tajbakhsh
Published: (2012)
by: Mahmood Tajbakhsh
Published: (2012)
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
Similar Items
-
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
by: Taghibakhshi, Ali, et al.
Published: (2026) -
BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
by: Venkatesh, Thejas, et al.
Published: (2026) -
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025) -
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
by: Taghibakhshi, Ali, et al.
Published: (2025) -
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
by: Taghibakhshi, Ali, et al.
Published: (2025)