In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Panigrahi, Abhishek, Liu, Bingbin, Malladi, Sadhika, Kakade, Sham, Goel, Surbhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive distillation induces an implicit curriculum
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
by: Sabbaghi, Mahdi, et al.
Published: (2024)
by: Sabbaghi, Mahdi, et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Conformal Language Model Reasoning with Coherent Factuality
by: Rubin-Toles, Maxon, et al.
Published: (2025)
by: Rubin-Toles, Maxon, et al.
Published: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
by: Lu, Taiming, et al.
Published: (2026)
by: Lu, Taiming, et al.
Published: (2026)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Knowledge Distillation with Training Wheels
by: Liu, Guanlin, et al.
Published: (2025)
by: Liu, Guanlin, et al.
Published: (2025)
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
by: Vujanic, Robin, et al.
Published: (2025)
by: Vujanic, Robin, et al.
Published: (2025)
Eliminating Position Bias of Language Models: A Mechanistic Approach
by: Wang, Ziqi, et al.
Published: (2024)
by: Wang, Ziqi, et al.
Published: (2024)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
by: Li, Chengye, et al.
Published: (2025)
by: Li, Chengye, et al.
Published: (2025)
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
by: Xue, Anton, et al.
Published: (2024)
by: Xue, Anton, et al.
Published: (2024)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
by: Hu, Yuezhou, et al.
Published: (2025)
by: Hu, Yuezhou, et al.
Published: (2025)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
by: Fang, Luyang, et al.
Published: (2025)
by: Fang, Luyang, et al.
Published: (2025)
A Study on the Calibration of In-context Learning
by: Zhang, Hanlin, et al.
Published: (2023)
by: Zhang, Hanlin, et al.
Published: (2023)
Sinkhorn Distance Minimization for Knowledge Distillation
by: Cui, Xiao, et al.
Published: (2024)
by: Cui, Xiao, et al.
Published: (2024)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
by: Shen, Ethan, et al.
Published: (2024)
by: Shen, Ethan, et al.
Published: (2024)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
by: Zhang, Hengyuan, et al.
Published: (2025)
by: Zhang, Hengyuan, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
LoCa: Logit Calibration for Knowledge Distillation
by: Yang, Runming, et al.
Published: (2024)
by: Yang, Runming, et al.
Published: (2024)
Confidence Preservation Property in Knowledge Distillation Abstractions
by: Vengertsev, Dmitry, et al.
Published: (2024)
by: Vengertsev, Dmitry, et al.
Published: (2024)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Similar Items
-
Progressive distillation induces an implicit curriculum
by: Panigrahi, Abhishek, et al.
Published: (2024) -
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022) -
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024) -
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)