Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bhardwaj, Kartikeya, Pandey, Nilesh Prasad, Priyadarshi, Sweta, Lee, Kyunggeun, Ma, Jun, Teague, Harris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
FouRA: Fourier Low Rank Adaptation
by: Borse, Shubhankar, et al.
Published: (2024)
by: Borse, Shubhankar, et al.
Published: (2024)
Video Reasoning without Training
by: Sridhar, Deepak, et al.
Published: (2025)
by: Sridhar, Deepak, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
by: Kong, Jason, et al.
Published: (2026)
by: Kong, Jason, et al.
Published: (2026)
Systems Toxicology of Bisphenol A: Mechanistic Overlap in Metabolic and Reproductive Disruption
by: Sweta Bhardwaj, et al.
Published: (2026)
by: Sweta Bhardwaj, et al.
Published: (2026)
QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design
by: Pandey, Nilesh Prasad, et al.
Published: (2026)
by: Pandey, Nilesh Prasad, et al.
Published: (2026)
Geometric Limits of Knowledge Distillation: A Minimum-Width Theorem via Superposition Theory
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
How Is Uncertainty Propagated in Knowledge Distillation?
by: Cui, Ziyao, et al.
Published: (2026)
by: Cui, Ziyao, et al.
Published: (2026)
Multi-aspect Knowledge Distillation with Large Language Model
by: Lee, Taegyeong, et al.
Published: (2025)
by: Lee, Taegyeong, et al.
Published: (2025)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
We Treat Children Not Guidelines
by: Aaron J. Stein, et al.
Published: (2025)
by: Aaron J. Stein, et al.
Published: (2025)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
Knowledge Distillation for Large Language Models
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
Self-Supervised Quantization-Aware Knowledge Distillation
by: Zhao, Kaiqi, et al.
Published: (2024)
by: Zhao, Kaiqi, et al.
Published: (2024)
Data-Augmented Quantization-Aware Knowledge Distillation
by: Kur, Justin, et al.
Published: (2025)
by: Kur, Justin, et al.
Published: (2025)
YSOs List IRAS 18456-0223
by: Pandey, Nilesh, et al.
Published: (2026)
by: Pandey, Nilesh, et al.
Published: (2026)
Stellar contents and Star Formation in IRAS 18456-0223
by: Pandey, Nilesh, et al.
Published: (2026)
by: Pandey, Nilesh, et al.
Published: (2026)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
by: Xin, Meng, et al.
Published: (2026)
by: Xin, Meng, et al.
Published: (2026)
Quantization dimensions for inhomogeneous bi-Lipschitz Iterated Function Systems
by: Priyadarshi, Amit, et al.
Published: (2023)
by: Priyadarshi, Amit, et al.
Published: (2023)
Quantization dimensions for the bi-Lipschitz recurrent Iterated function systems
by: Priyadarshi, Amit, et al.
Published: (2022)
by: Priyadarshi, Amit, et al.
Published: (2022)
Fast and Accurate Contextual Knowledge Extraction Using Cascading Language Model Chains and Candidate Answers
by: Harris, Lee
Published: (2025)
by: Harris, Lee
Published: (2025)
Interpretable Question Answering with Knowledge Graphs
by: Aneja, Kartikeya, et al.
Published: (2025)
by: Aneja, Kartikeya, et al.
Published: (2025)
Delta Knowledge Distillation for Large Language Models
by: Cao, Yihan, et al.
Published: (2025)
by: Cao, Yihan, et al.
Published: (2025)
Quantum Knowledge Distillation for Large Language Models
by: Li, Lingxiao, et al.
Published: (2025)
by: Li, Lingxiao, et al.
Published: (2025)
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
DistillER: Knowledge Distillation in Entity Resolution with Large Language Models
by: Zeakis, Alexandros, et al.
Published: (2026)
by: Zeakis, Alexandros, et al.
Published: (2026)
Harnessing Chaotic Signals for Wireless Information and Power Transfer
by: Mukherjee, Priyadarshi, et al.
Published: (2025)
by: Mukherjee, Priyadarshi, et al.
Published: (2025)
Advanced Knowledge Transfer: Refined Feature Distillation for Zero-Shot Quantization in Edge Computing
by: Hong, Inpyo, et al.
Published: (2024)
by: Hong, Inpyo, et al.
Published: (2024)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
by: Fang, Luyang, et al.
Published: (2025)
by: Fang, Luyang, et al.
Published: (2025)
Chaotic Waveform-based Signal Design for Noncoherent SWIPT Receivers
by: Mukherjee, Priyadarshi, et al.
Published: (2023)
by: Mukherjee, Priyadarshi, et al.
Published: (2023)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
by: Zhou, Chenxi, et al.
Published: (2025)
by: Zhou, Chenxi, et al.
Published: (2025)
Oh lucky man
Published: (1998)
Published: (1998)
Oh, That Library Loan
by: Shollenberger, Richard C.
Published: (1972)
by: Shollenberger, Richard C.
Published: (1972)
Rats! Oh No, Not Rats!
by: Strong, Gary E.
Published: (1987)
by: Strong, Gary E.
Published: (1987)
We Need Knowledge Distillation for Solving Math Word Problems
by: Shen, Zhenquan, et al.
Published: (2025)
by: Shen, Zhenquan, et al.
Published: (2025)
Knowledge Distillation for Temporal Knowledge Graph Reasoning with Large Language Models
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
PQV-Mobile: A Combined Pruning and Quantization Toolkit to Optimize Vision Transformers for Mobile Applications
by: Bhardwaj, Kshitij
Published: (2024)
by: Bhardwaj, Kshitij
Published: (2024)
Knowledge Distillation of Black-Box Large Language Models
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Similar Items
-
Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024) -
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024) -
FouRA: Fourier Low Rank Adaptation
by: Borse, Shubhankar, et al.
Published: (2024) -
Video Reasoning without Training
by: Sridhar, Deepak, et al.
Published: (2025) -
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)