Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Binici, Kuluhan, Wu, Weiming, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay
by: Binici, Kuluhan, et al.
Published: (2022)
by: Binici, Kuluhan, et al.
Published: (2022)
Condensed Data Expansion Using Model Inversion for Knowledge Distillation
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Split Knowledge Distillation for Large Models in IoT: Architecture, Challenges, and Solutions
by: Li, Zuguang, et al.
Published: (2024)
by: Li, Zuguang, et al.
Published: (2024)
Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher
by: Ji, Guangda, et al.
Published: (2020)
by: Ji, Guangda, et al.
Published: (2020)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
From Teacher to Student: Tracking Memorization Through Model Distillation
by: Singh, Simardeep
Published: (2025)
by: Singh, Simardeep
Published: (2025)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026)
by: Fang, Luyang, et al.
Published: (2026)
Teaching the Teacher: The Role of Teacher-Student Smoothness Alignment in Genetic Programming-based Symbolic Distillation
by: Dhar, Soumyadeep, et al.
Published: (2025)
by: Dhar, Soumyadeep, et al.
Published: (2025)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
by: Li, Chao, et al.
Published: (2025)
by: Li, Chao, et al.
Published: (2025)
Efficient and Robust Knowledge Distillation from A Stronger Teacher Based on Correlation Matching
by: Niu, Wenqi, et al.
Published: (2024)
by: Niu, Wenqi, et al.
Published: (2024)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
A Functional Perspective on Knowledge Distillation in Neural Networks
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
Online Adversarial Knowledge Distillation for Graph Neural Networks
by: Wang, Can, et al.
Published: (2021)
by: Wang, Can, et al.
Published: (2021)
Learning Privacy-Preserving Student Networks via Discriminative-Generative Distillation
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer
by: Udayangani, Nilushika, et al.
Published: (2026)
by: Udayangani, Nilushika, et al.
Published: (2026)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
by: Kim, Gyeongman, et al.
Published: (2024)
by: Kim, Gyeongman, et al.
Published: (2024)
Knowledge Distillation on Spatial-Temporal Graph Convolutional Network for Traffic Prediction
by: Izadi, Mohammad, et al.
Published: (2024)
by: Izadi, Mohammad, et al.
Published: (2024)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
by: Shin, Hyunjune, et al.
Published: (2024)
by: Shin, Hyunjune, et al.
Published: (2024)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
DNAD: Differentiable Neural Architecture Distillation
by: Rao, Xuan, et al.
Published: (2025)
by: Rao, Xuan, et al.
Published: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
by: Xiang, Qianlong, et al.
Published: (2024)
by: Xiang, Qianlong, et al.
Published: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
Linear Projections of Teacher Embeddings for Few-Class Distillation
by: Loo, Noel, et al.
Published: (2024)
by: Loo, Noel, et al.
Published: (2024)
DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
by: Wang, Zexin, et al.
Published: (2025)
by: Wang, Zexin, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Towards Mitigating Architecture Overfitting on Distilled Datasets
by: Zhong, Xuyang, et al.
Published: (2023)
by: Zhong, Xuyang, et al.
Published: (2023)
Graph Knowledge Distillation to Mixture of Experts
by: Rumiantsev, Pavel, et al.
Published: (2024)
by: Rumiantsev, Pavel, et al.
Published: (2024)
Not All Timesteps Matter Equally: Selective Alignment Knowledge Distillation for Spiking Neural Networks
by: Sun, Kai, et al.
Published: (2026)
by: Sun, Kai, et al.
Published: (2026)
Dynamic Temperature Scheduler for Knowledge Distillation
by: Islam, Sibgat Ul, et al.
Published: (2025)
by: Islam, Sibgat Ul, et al.
Published: (2025)
Membership and Memorization in LLM Knowledge Distillation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
Similar Items
-
Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay
by: Binici, Kuluhan, et al.
Published: (2022) -
Condensed Data Expansion Using Model Inversion for Knowledge Distillation
by: Binici, Kuluhan, et al.
Published: (2024) -
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
by: Aggarwal, Shivam, et al.
Published: (2023) -
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022) -
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
by: Nagar, Aishik, et al.
Published: (2024)