Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Tong, Liu, Long, Hu, Yihang, Chen, Hu, Chen, Shifeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation
by: Zhang, Zhe, et al.
Published: (2026)
by: Zhang, Zhe, et al.
Published: (2026)
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
by: Hu, Chengming, et al.
Published: (2023)
by: Hu, Chengming, et al.
Published: (2023)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026)
by: Fang, Luyang, et al.
Published: (2026)
Toward Fair Graph Neural Networks Via Dual-Teacher Knowledge Distillation
by: Li, Chengyu, et al.
Published: (2024)
by: Li, Chengyu, et al.
Published: (2024)
Towards the Law of Capacity Gap in Distilling Language Models
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
by: Aslam, Muhammad Haseeb, et al.
Published: (2025)
by: Aslam, Muhammad Haseeb, et al.
Published: (2025)
Exploring Dark Knowledge under Various Teacher Capacities and Addressing Capacity Mismatch
by: Fan, Wen-Shu, et al.
Published: (2024)
by: Fan, Wen-Shu, et al.
Published: (2024)
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
by: Hu, Zhengyu, et al.
Published: (2026)
by: Hu, Zhengyu, et al.
Published: (2026)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
by: Li, Chao, et al.
Published: (2025)
by: Li, Chao, et al.
Published: (2025)
Student Capacity Moderates Knowledge Distillation Effectiveness: A Systematic Study Across ResNet Teacher-Student Pairs on CIFAR-10
by: Yasar, Umut Onur
Published: (2026)
by: Yasar, Umut Onur
Published: (2026)
Distilling Realizable Students from Unrealizable Teachers
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Adversarial Dual On-Policy Distillation from Expressive Teacher
by: Wan, Zhenglin, et al.
Published: (2026)
by: Wan, Zhenglin, et al.
Published: (2026)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
How to Train the Teacher Model for Effective Knowledge Distillation
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
Embedding Compression for Teacher-to-Student Knowledge Transfer
by: Ding, Yiwei, et al.
Published: (2024)
by: Ding, Yiwei, et al.
Published: (2024)
Linear Projections of Teacher Embeddings for Few-Class Distillation
by: Loo, Noel, et al.
Published: (2024)
by: Loo, Noel, et al.
Published: (2024)
Continual Policy Distillation from Distributed Reinforcement Learning Teachers
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines
by: Morad, Itai, et al.
Published: (2026)
by: Morad, Itai, et al.
Published: (2026)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
Teaching the Teacher: The Role of Teacher-Student Smoothness Alignment in Genetic Programming-based Symbolic Distillation
by: Dhar, Soumyadeep, et al.
Published: (2025)
by: Dhar, Soumyadeep, et al.
Published: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
Knowledge Distillation Based on Transformed Teacher Matching
by: Zheng, Kaixiang, et al.
Published: (2024)
by: Zheng, Kaixiang, et al.
Published: (2024)
From Teacher to Student: Tracking Memorization Through Model Distillation
by: Singh, Simardeep
Published: (2025)
by: Singh, Simardeep
Published: (2025)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
by: Shin, Hyunjune, et al.
Published: (2024)
by: Shin, Hyunjune, et al.
Published: (2024)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Class-wise Federated Unlearning: Harnessing Active Forgetting with Teacher-Student Memory Generation
by: Li, Yuyuan, et al.
Published: (2023)
by: Li, Yuyuan, et al.
Published: (2023)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
by: Awal, Md. Abdul, et al.
Published: (2025)
by: Awal, Md. Abdul, et al.
Published: (2025)
Bridging the Gap Between Average and Discounted TD Learning
by: Tian, Haoxing, et al.
Published: (2026)
by: Tian, Haoxing, et al.
Published: (2026)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
by: Lu, Taiming, et al.
Published: (2026)
by: Lu, Taiming, et al.
Published: (2026)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
by: Xia, Mingxuan, et al.
Published: (2025)
by: Xia, Mingxuan, et al.
Published: (2025)
SFedKD: Sequential Federated Learning with Discrepancy-Aware Multi-Teacher Knowledge Distillation
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
by: Khani, Nikhil, et al.
Published: (2024)
by: Khani, Nikhil, et al.
Published: (2024)
Bridging the Domain Gap in Equation Distillation with Reinforcement Feedback
by: Ying, Wangyang, et al.
Published: (2025)
by: Ying, Wangyang, et al.
Published: (2025)
Similar Items
-
Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation
by: Zhang, Zhe, et al.
Published: (2026) -
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
by: Hu, Chengming, et al.
Published: (2023) -
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
by: Wu, Lirong, et al.
Published: (2024) -
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022) -
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026)